Plasma Cell-free DNA 5-Hydroxymethylcytosine and Whole-Genome Sequencing Signatures for Early Detection of Esophageal Cancer

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

Plasma cfDNA 5hmC signatures and low-pass whole-genome sequencing improved early esophageal cancer detection, particularly for Stage 0.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This preprint evaluated whether plasma cell-free DNA (cfDNA) 5-hydroxymethylcytosine (5hmC) signatures, alone and combined with low-pass whole-genome sequencing (WGS) features, can detect early esophageal squamous cell carcinoma (ESCC). The authors profiled 5hmC using 100 ESCC patients and 71 healthy controls from a Southern China cohort, built a diagnostic model from 273 5hmC features, and validated it internally and against an external Northern cohort (150 ESCC, 183 controls) with an explicitly reported caveat that performance was low for Stage 0 (accuracy 33.3%) despite higher accuracy in Stages I–IV. They found conserved cfDNA 5hmC motif changes across cohorts, and that adding low-pass WGS improved overall discrimination (AUC up to 0.934), with notably better Stage 0 accuracy (up to 80%). The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background: Esophageal cancer is one of globally high incidence and mortality disease. Its early stage has rarely obvious symptoms. Compared to conventional endoscopy diagnosis, liquid biopsy is an emerging non-invasive method for cancer early detection. Methods We enrolled 100 esophageal squamous cell carcinoma (ESCC) patients and 71 healthy individuals as a Southern China cohort and performed 5-hydroxymethylcytosine (5hmC) sequencing on their plasma cell-free DNA (cfDNA). A Northern cohort of cfDNA 5hmC dataset with 150 ESCC patients and 183 health individuals were downloaded for validation. A diagnostic model was firstly developed based on cfDNA 5hmC signatures and then improved by low-pass whole genome sequencing (WGS) features of cfDNA. Results Conserved cfDNA 5hmC modification motifs were observed in the two independent ESCC cohorts. A diagnostic model with 273 5hmC features based on randomly-selected two-thirds samples of the Southern China cohort was validated independently in the left one-thirds samples and the whole Northern China cohort, achieved an AUC of 0.810 and 0.862 with sensitivities of 69.3–74.3% and specificities of 82.4–90.7%, respectively. The performance was well maintained in Stage I to Stage IV, with accuracy of 70%-100%, but low in Stage 0, with accuracy of 33.3%. Low-pass WGS of cfDNA improved the AUC to 0.934 with a sensitivity of 82.4%, a specificity of 88.2% and an accuracy of 84.3%, particularly significantly in Stage 0 with the accuracy up to 80%. Conclusions This study suggests that the blood-based 5hmC integrated with low-pass WGS model could improve the accurate diagnosis of early stage ESCC, particularly for very early stage ESCC.
Full text 170,102 characters · extracted from preprint-html · click to expand
Plasma Cell-free DNA 5-Hydroxymethylcytosine and Whole-Genome Sequencing Signatures for Early Detection of Esophageal Cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Plasma Cell-free DNA 5-Hydroxymethylcytosine and Whole-Genome Sequencing Signatures for Early Detection of Esophageal Cancer Di Lu, Xuanzhen Wu, Shuangxiu Wu, Hui Li, Xuebin Yan, Jianxue Zhai, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1375061/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Esophageal cancer is one of globally high incidence and mortality disease. Its early stage has rarely obvious symptoms. Compared to conventional endoscopy diagnosis, liquid biopsy is an emerging non-invasive method for cancer early detection. Methods We enrolled 100 esophageal squamous cell carcinoma (ESCC) patients and 71 healthy individuals as a Southern China cohort and performed 5-hydroxymethylcytosine (5hmC) sequencing on their plasma cell-free DNA (cfDNA). A Northern cohort of cfDNA 5hmC dataset with 150 ESCC patients and 183 health individuals were downloaded for validation. A diagnostic model was firstly developed based on cfDNA 5hmC signatures and then improved by low-pass whole genome sequencing (WGS) features of cfDNA. Results Conserved cfDNA 5hmC modification motifs were observed in the two independent ESCC cohorts. A diagnostic model with 273 5hmC features based on randomly-selected two-thirds samples of the Southern China cohort was validated independently in the left one-thirds samples and the whole Northern China cohort, achieved an AUC of 0.810 and 0.862 with sensitivities of 69.3–74.3% and specificities of 82.4–90.7%, respectively. The performance was well maintained in Stage I to Stage IV, with accuracy of 70%-100%, but low in Stage 0, with accuracy of 33.3%. Low-pass WGS of cfDNA improved the AUC to 0.934 with a sensitivity of 82.4%, a specificity of 88.2% and an accuracy of 84.3%, particularly significantly in Stage 0 with the accuracy up to 80%. Conclusions This study suggests that the blood-based 5hmC integrated with low-pass WGS model could improve the accurate diagnosis of early stage ESCC, particularly for very early stage ESCC. esophageal cancer early diagnosis 5-hydroxymethylcytosine low-passwhole-genome sequencing cell-free DNA Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Esophageal cancer is a global problem threatening people’s health and declining people’s life expectancy which ranks the seventh in terms of incidence and the sixth in mortality worldwide in 2020 with two main subtypes: esophageal squamous cell carcinoma (ESCC) and esophageal adenocarcinoma (EAC)[ 1 ].There are distinct geographic variations between the two main subgroups of esophageal cancer. The incidence of ESCC is apparently higher in China, Central Asia, East and South of Africa while most of the EAC cases occur in patients living in high income countries like North America and Europe[ 2 ].Most of the patients were already suffered from ESCC with an advanced stage when they sought medical advices as early esophageal cancer was rarely accompanied by distinct symptoms, thus, neither the surgical procedures nor adjuvant therapies could lead to satisfied outcomes[ 3 ]. So, an urgent demand for a new method to help making early diagnosis of esophageal carcinoma is desired. Generally, tissue biopsy like endoscopy is the most widespread procedure used to make detection of esophageal cancer and help formulating stages for further therapeutic schedule. However, there remains obvious shortcomings and inadequacies like inferior diagnostic efficiency, uncomfortable checking process and expensive spending which may lead to poor adherence to subsequent treatments in return[ 4 , 5 ]. Compared to conventional tissue biopsy, liquid biopsy is an emerging non-invasive method to make early detection which can also make further molecular and genic analysis using these liquid specimens[ 6 ]. Thereinto, cell-free DNA (cfDNA) is deliberated from the cell existing in blood, lymph, bile, milk, urine, saliva, mucous suspension, spinal fluid and amniotic fluid originating from apoptosis, necrosis, phagocytosis, nocosis and active secretion[ 7 ], and was proved to be a approving biomarker to make screening, detection and monitoring of cancers[ 8 ]. However, there remains some challenges in the application of cfDNA for early detection of cancers like hyposecretion, variability and polyphyly[ 9 ]. The efficiency of DNA methylation as an epigenetic mark has been previously validated[ 10 ], furthermore, the association between DNA methylation and tumorigenesis has been verified too[ 11 , 12 ]. As the conversion of cytosine, 5-hydroxymethylcytosine (5hmC) is oxidized from 5-methylcytosine(5mC) by the enzyme ten-eleven translocation 1(TET1) which is one of the three enzymes of TET(ten-eleven translocation) family and is recognized as a better marker to detect gene expression[ 13 ]. Although there are fewer 5hmC existing in tissues and more challenges to accomplish precise detection, it has been proved that 5hmC has a superior permissive effect on gene expression and exhibits more tissue specificity. Besides, plenty of studies have reported that 5hmC in cfDNA was used as a promising marker to detect cancers like early-stage pancreatic cancer, non-small-cell lung cancer (NSCLC), hepatocellular carcinoma (HCC), blood and colon cancer[ 14 – 19 ]. In addition, a recent study showed ctDNA in plasma had distinct features, such as shorter fragment size, special motif and nucleosome footprint (NF) in comparison with cfDNA of healthy people[ 20 ], which could be applied to an advanced technology (low-pass whole-genome sequencing (WGS)) to get more accurate outcomes in identifying ctDNA from plasma. WGS of plasma cfDNA was a new empowering technology used for non-invasion low-disease-burden diagnosis. This technology provided ultra-sensitive detection of ctDNA from cfDNA of healthy individuals via genome-wide mutational integration[ 20 ]. Recent studies proved its good performance on early diagnosis of lung cancer[ 21 , 22 ], hepatocellular carcinoma (HCC)[ 23 ] and malignant peripheral nerve sheath tumor (MPNST)[ 24 ]. Therefore, employing 5hmC in plasma cfDNA as a biomarker to distinguish the differences between ESCC patients and healthy control people (HC) is worth trying, and we aimed to determine 5hmC signatures in cfDNA using machine learning methods as early diagnostic biomarkers for ECSS in a prospective cohort study. The low-pass WGS technology was also attempted as a complementary technique for identification of early ESCC from healthy people. Results Samples composition and study design A total of 150 adult subjects were prospectively enrolled in this study from March 2019 to December 2020, including patients with Early ESCC (stages 0, IA and IB, n = 50), and middle and advanced (Mid-Ad) ESCC (stages II, III and IVA, n = 50) and healthy control individuals (HC, n = 50) from Nanfang Hospital of Southern Medical University, as a Southern ESCC cohort since the patients were mainly from the South of China. Detailed clinical information regarding gender, age, BMI, living habits, differentiation of cancer, TNM stages and surgery selections were illustrated in Table 1 . And after conducting a brief statistical analysis using the data of the 150 patients collected from our hospital, we observed that there were no apparent differences among most of the variates like gender, age, BMI, habits and the incidences of hypertension and diabetes ( P > 0.05), except for family history ( P < 0.05), which suggested that family history might play an important role in tumorigenesis. Table 1 Summary of demographic and clinicopathological characteristics of all the participants in this study Characteristics HC (n = 50) Early ESCC (n = 50) Mid-Ad ESCC (n = 50) P value Demographic Gender Male 37 37 37 Female 13 13 13 Age (years) (mean, range) 58 (50–70) 60 (50–70) 59 (50–70) 0.15 BMI (mean, range) 22.9 (15.2–28.7) 22.2 (17.6–29.8) 22.0 (16.4–27.3) 0.28 Salted foods (like/dislike) 13 (29)* 16 (34) 25 (25) 0.106 Smoking (yes/no) 22 (20)* 30 (20) 31 (19) 0.667 Drinking (yes/no) 20 (22)* 20 (30) 29 (21) 0.201 Fresh vegetables and fruits (like/dislike) 30 (12)* 38 (12) 34 (16) 0.654 Family history (with/without) 0 (42)* 10 (40) 6 (44) 0.005 Hypertension (with/without) 7 (35)* 13 (37) 7 (43) 0.282 Diabetes (with/without) 1 (41)* 6 (44) 5 (45) 0.212 Clinical Differentiation G0 NA 27 NA G1 NA 16 20 G2 NA 6 21 G3 NA 1 9 TNM stages 0 NA 27 NA IA NA 3 NA IB NA 20 NA IIA NA NA 19 IIB NA NA 7 IIIA NA NA 1 IIIB NA NA 19 IVA NA NA 4 Surgery MATHE NA 12 NA McKeown NA 9 24 Sweet NA 18 25 IvorLewis NA NA 1 ESD NA 11 NA Note: Patients were classified into 3 groups; HC, healthy controls; Mid-Ad, middle-advanced; ESCC, esophageal squamous cell carcinoma; MATHE: mediastinoscope-assisted transhiatal esophagectomy; ESD: endoscopic submucosal dissection; P value in chi-square test or t test. *There were 8 unknown findings. Additional 21 HC samples were acquired from a previously published study[ 25 ]. The HC cohort was identified as satisfying the study inclusion criteria, which were specifically negative for any form of cancer. Detailed pathological and clinical information was obtained from clinical records and illustrated in Table 1 and supplementary Table S1. The subject age and gender distributions among HC, Early ESCC and Mid-Ad ESCC groups were unbiased based on the Kruskal-Wallis H-test ( P value = 0.11, supplementary Figure S1A) and Fisher test ( P value > 0.5, supplementary Figure S1B), respectively. Our primary aim was to develop a convenient and excellent diagnostic model using new biotechnology reflecting genome-wide signatures from plasma cfDNA to distinguish patients with Early ESCC from HC subjects. For biomarker screening and classifier model construction for early diagnosis of ESCC, the ESCC patients and HC individuals were randomly divided into two groups: about 2/3 individuals as a training set and 1/3 individuals as an independent test set (also named as the internal test set) (Fig. 1 and supplementary Figure S2). We also downloaded 5hmC-sequencing data from previously published studies of ESCC (as another independent external test set, also named as a Northern ESCC cohort since these ESCC patients were mainly from the North of China)[ 26 ]. Conserved 5hmc Modification Changes In Escc, Potential Biomarkers For Diagnosis To explore the distribution patterns of hydroxymethylation in plasma cfDNAs across the genome, we first identified 5hmC-enriched regions in each sample and subsequently determined as reliable peaks by high enrichment and significance as described in Methods. Comparison of global 5hmC modification levels between ESCC and HC groups revealed significant differences. The density of 5hmC peaks number were calculated and exhibited a broader distribution in ESCC group compared with the HC group that displayed a sharper and narrower curve (Fig. 2 A). 5hmC peaks number in ESCC cfDNAs cohort was significantly more than that in HC cohort (Mann-Whitney U-test, P value = 3.11×10 − 5 , supplementary Figure S3) and showed a gradually increasing trend from stage 0 to stage IV patients (Mann-Kendall Test, P value = 1.65 ×10 − 3 , Fig. 2 B), indicating that 5hmC modification changes might play an important role in promoting tumorigenesis and progression of esophagus cancer. In addition, increased 5hmC modification levels within promoter and genebody regions in ESCC group were observed by 5hmC metagene profile analysis (Fig. 2 C), consistent with previous observation in ESCC[ 26 ]. To further compare the cfDNA 5hmC modification difference between the two groups, we determined differential 5hmC peaks with DESeq2 package ( P value < 0.01) and identified 398 5hmC up-regulation peaks and 227 5hmC down-regulation peaks in ESCC groups by comparing with HC groups. 5hmC up-regulation peaks from ESCC patients were significant enriched in promoters (28.14%) and 1st intron regions (15.58%) (i.e., mainly in the regulation regions of a gene) on the whole genome level, while more 5hmC down-regulation peaks were located in other introns (37.44%) and distal intergenic regions (39.21%) (Fig. 2 D). Next, 5hmC motif enrichment analysis was performed in up-regulation and down-regulation regions of 5hmC signals in ESCC samples to further understand the correlation of 5hmC changes with potential interactions of binding proteins. The 5hmC up-regulation peaks in ESCC was significantly enriched in ERG motif ( P = 1e-5, 28.83%), followed by ETS1 ( P = 1e-4, 20.25%) and ETV2 motif ( P = 1e-3, 17.79%), all of which belong to the large family of ETS transcription factors and bind to the consensus DNA sequence 5'-AGGAA-3' (left in Fig. 2 E), most of which are downstream nuclear targets of Ras-MAP kinase signaling, and associate with cell development, differentiation, proliferation, apoptosis and tissue remodeling. The deregulation of ETS genes results in the malignant transformation of cells and is often observed in various types of malignant tumors[ 27 ]. In contrast, the motif of GATA3 ( P = 1e-5, 29.23%), GATA4 ( P = 1e-5, 22.31%) and TRPS1 ( P = 1e-4, 33.85%) were observed in 5hmC down-regulation peaks (right in Fig. 2 E). GATA transcription factor mutations or dysregulated GATA expression were reported to associate with the wide range of diseases and pathologic phenotypes[ 28 ]. These are consistent with previously published study showing that ETS motif and GATA motif were identified in 5hmC-gain and 5hmC-loss regions for esophageal cancer, respectively[ 26 ]. On the other hand, we further compared the enriched 5hmC motifs among our Southern ESCC cohort, Northern ESCC cohort[ 26 ] and colorectal cancer cohort[ 25 ] and observed. For the top 10-significantly hit 5hmC motifs, about 30.0%-50.0% 5hmC motifs shared between the Southern and the Northern ESCC cohorts, obviously more than those shared with CRC patients (about 10.0%-20.0%) (Fig. 2 F, supplementary Table S2). These results showed that the unique signature of plasma cfDNA 5hmC may represents a stable biomarker for discriminating ESCC and healthy individuals and could be potentially useful for early ESCC diagnosis. Screening, Validation And Performance Of Candidate 5hmc Biomarkers And Classifier Since the results of 5hmC up-regulation peaks were significantly enriched in promoter and 1st intron regions and 5hmC down-regulation peaks were significantly enriched in other introns and distal intergenic regions in the ESCC patients, which indicated that differential 5hmC signals in promoter and certain genebody regions might be utilized for identifying candidate 5hmC biomarkers to develop a convenient and stable diagnostic model. All 171 samples were randomly separated into two groups (54 HC, 34 Early ESCC and 31 Mid-Ad ESCC for training set, the remaining samples for internal test set) for development and evaluation of diagnostic model. The trained model was also validated on the independent previously-published Northern ESCC dataset (external test set) included 150 esophageal cancer and 183 HC samples[ 26 ], which was profiled for 5hmC using a nano-hmC-Seal method similar to the one used in our study. Firstly, 925 candidate 5hmC marker genes derived from promoter and genebody regions were selected by Wilcoxon rank-sum test P values < 0.001 in the training set. Using Recursive feature elimination - Cross Validation (RFECV) approach, we further identified a disease-specific panel of 273 5hmC marker genes according to the contribution of each gene in the model (supplementary Table S3), and the distinct 5hmC landscapes in cfDNA showed apparent separation between ESCC and HC groups by unsupervised hierarchical clustering analysis (Fig. 3 A). Similarly, the principal component analysis (PCA) also demonstrated distinct signatures that could discriminate the majority of the individuals of the two groups (Fig. 3 B). To explore the biological significance of 273 marker genes with differential 5hmC signals, we performed gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analysis and found that 273 5hmC biomarkers were enriched in pathways associated with cancer and metastasis and mapped to tumor-related genes (Fig. 3 C). For instance, Fig. 3 D exhibited the IGV plot of the high-weight biomarker located at the FOXK1 gene, which plays an oncogenic role in the development of esophageal cancer[ 29 ]. The ESCC samples showed increased 5hmC modification level at FOXK1 gene when compared with the healthy controls. The classifier model based on 273 marker gene illustrated decent capacity for distinguishing ESCC patients from HC individuals in both internal test set Area under curve (AUC) = 0.810 (95% CI: 0.693–0.927); sensitivity = 74.3%; specificity = 82.4%) and external test set (AUC = 0.862 (95% CI: 0.822–0.902); sensitivity = 69.3%; specificity = 90.7%) (Fig. 3 E and 3 F). The prediction performance of classifier model in external test set was superior to that of internal test set, probably because of patients with 0 stage accounting for 27% (27/100) of our ESCC cohort who might be misclassified as HC individuals. We next assess the prediction accuracy of 5hmC biomarker classifier for different clinical stages in our ESCC patient subgroups. Box plots exhibited that the probability of being predicted as cancer gradually increasing with the progression of cancer stage in the test set (Fig. 3 G). The 5hmC score between Early ESCC (stage 0 and I) and HC individuals demonstrated statistically significant disparity ( P value = 4.35×10 − 2 , supplementary Figure S4), which suggested the capacity of the 5hmC model to discriminate Early ESCC patients from HC individuals. Meanwhile, 5hmC score also accurately distinguished between stage I and HC samples ( P value = 4.54×10 − 3 ), stage 0 and stage I, II, III-IV ( P values = 4.64×10 − 2 , 1.46×10 − 2 and 9.85×10 − 3 , respectively, Fig. 3 G). However, in differentiating stage 0 from HC samples, 5hmC score was not the best diagnostic feature ( P value = 0.40, Fig. 3 G), and the diagnostic accuracy (33.3%, supplementary Figure S5B) need to be further improved. Integrated model based on cfDNA signatures of low-pass WGS and 5hmC biomarkers improved diagnostic score for Early ESCC To explore the prediction potential of plasma cfDNA and search for more effective biomarkers, we employed low-pass whole-genome sequencing (WGS) to acquire genome-wide 5’ end motif[ 30 ], nucleosome footprint (NF)[ 31 ] and fragmentation[ 8 ] profiles from 71 HC and 93 ESCC samples with enough plasma cfDNA. Differential 5’ end motif was identified by Wilcoxon rank-sum test P value < 0.001 and exhibited apparent separation for ESCC against HC samples by unsupervised hierarchical clustering analysis (Fig. 4 A). Meanwhile, NF heatmap analysis indicated that genes with differential reads coverage between promoter and background regions (Wilcoxon rank-sum test, P value < 0.001) held power to distinguish ESCC from HC (Fig. 4 B). Additionally, our data detected that the cfDNA fragment size of ESCC was more variable and much shorter (median size < 150 base-pairs (bp)) than that of HC group (Fig. 4 C). Collectively, all three genome features of cfDNA showed promising diagnostic potential for ESCC. As illustrated in Fig. 1 and supplementary Figure S2, healthy control individuals and patients with Early ESCC and Mid-Ad ESCC were randomly assigned into a training set (about 2/3 of samples, including 54 HC, 30 Early ESCC and 29 Mid-Ad ESCC) and a test set (about 1/3 of samples, including 17 HC, 15 Early ESCC and 19 Mid-Ad ESCC). Candidate biomarkers of the above three genomic features (5’ end motif, NF and fragmentation) were screened to distinguish ESCC and HC in the training set. Least absolute shrinkage and selection operator (LASSO) regression method were applied to further reduce the number of candidate biomarkers. Eventually, 120 differential motif types, 170 differential NF genes and 10 fragment areas were selected for further analysis (supplementary Table S4A-S4C). Then, the Support Vector Machine (SVM) method was implemented for individual genomic feature-based model construction. The set of 120 motif types resulted in a discrimination model achieving an AUC value of 0.870 (95% CI: 0.769–0.972) with sensitivity of 73.5% at specificity of 82.4% for ESCC patient classification in the test set (Fig. 4 D, supplementary Figure S5A). The motif prediction model exhibited superior diagnostic power outperformed that of NF and fragmentation classifier model in cfDNA, which achieving an inferior performance with an AUC value of 0.813 (95% CI: 0.665–0.961) and 0.806 (95% CI: 0.677–0.936), respectively (Fig. 4 D). Compared with the 5’ end motif model, the fragmentation model demonstrated a higher sensitivity of 79.4% at the same specificity of 82.4%, and NF model showed excellent sensitivity of 91.2% but a lower specificity of 70.6% (supplementary Figure S5A). Because the above three genomic features and 5hmC depict different aspects of the genome, we envisioned that conjoint analysis would improve diagnostic power. Using the predictive score of the four individual models as input features, an integrated diagnostic model of combining low-pass WGS cfDNA signatures and 5hmC biomarkers was constructed. The diagnostic power of integrated model outperformed any individual genomic features, achieving an excellent AUC value of 0.934 (95% CI: 0.867-1.000) with a sensitivity of 82.4%, specificity of 88.2% and accuracy of 84.3% for ESCC patient classification in the test set (Fig. 4 D, 4 E). The combine scores showed an increasing trend from HC to ESCC, noting the significantly higher scores in ESCC patients with stage 0 and stage I than in subjects of HC in the test set ( P values = 1.10×10 − 2 and 3.31×10 − 5 , respectively, Fig. 4 F), reinforcing the idea that integrated model had great potentials as a new strategy for ESCC early diagnosis and surveillance. The integrated classifier model had good but slightly reduced power to call stage III-IV patients (Fig. 4 F), who in general suffered from metastasis to various tissues and are expected to have more complex tumorous DNA profiles. In addition, we compared the diagnostic performance between 5hmC model and integrated model on different clinical stage of ESCC samples in test set (Fig. 4 G). For Early ESCC patients, especially patients with stage 0 who have low-disease burden, the performance of integrated model was obviously superior to that of 5hmC model (Fig. 4 G), showing evidently higher prediction accuracy (80.0% vs 33.3%) (supplementary Figure S5B). Integrated model and 5hmC model validated similar performance in stage II patients. Instead, for advanced ESCC patients (stage IIIB and IVA), 5hmC model showed an improved diagnostic performance compared to integrated model. These data demonstrated that genome-wide integration was a sensitive and robust approach which could perform better than 5hmC-based methods for accurate diagnosis of early stage esophageal carcinoma. Discussion We ultimately constructed an integrated classifier using low-pass WGS and 5hmC biomarkers to realize our primal goal for early detection of esophageal cancer which may bring convenience to clinical practice. At first step, an investigation of the enriched 5hmC motifs in genome was conducted, and the outcomes showed that these existing some enriched 5hmC motifs in different regions like ERG, ETS1 and ETV2 which were proved to associate with cell developments, differentiation, proliferation, apoptosis and tissue remodeling, as well as GATA3, GATA4 and TRPS1 had been reported to be connected with a wide range of diseases and pathologic phenotypes, and both of these findings were consistent with our previous awareness about the biological behavior of tumors. We then compared the enriched 5hmC motifs among our Southern ESCC cohort (with 150 pb paired-end sequencing mode), the Northern ESCC cohort (with 50 bp paired-end sequencing mode)[ 26 ] and colorectal cancer cohort (with 150 pb paired-end sequencing mode)[ 25 ] which suggested that some 5hmC motif signal changes were more conservative in ESCC patients than in other patients (such as CRC), which were not effected by different sequencing platforms. These ESCC-unique 5hmC motifs mainly consisted of up-regulated ERG, ETS1 and ETV2 of ETS family, which were reported as cancer-associated transcription factors regulating Ras-MAP kinase signaling pathways[ 27 , 32 – 37 ]. Up-regulated EHF and ELF3 of ETS family uniquely consisted in the Northern ESCC 5hmC motifs, indicating the bias possibly due to geographic distribution or different pathological subtype compositions between these two ESCC cohorts. Other functions, such as vasculature development, muscle structure development, cell part morphogenesis/development and hippo signaling pathway were also consistent with those reported in the Northern ESCC cohort, indicating their essential roles in ESCC tumorigenesis and developments. In addition, other different functions among 5hmC biomarkers between our study and the Northern ESCC cohort might reveal different molecular mechanisms in different pathological subgroups of ESCC cohort or different epigenetic modifications resulting from different living habitats, which needed further investigation on large scale population studies. Recently, some relevant studies showed somewhat dissatisfactory early detection efficiency which results were around 50% and 62.5% in sensitivity on esophageal cancer stage 0 and I respectively using cfDNA methylation method[ 38 ], and some of the others exhibited a good sensitivity of 93.75% and specificity of 85.71% (AUC = 0.972) in distinguishing ESCC from HC groups but slightly ignored the discussion focusing on the early stage using 5hmC method[ 26 ]. In the present study, using our selected 273 5hmC marker genes, we indeed succeeded in distinguishing ESCC patients from HC individuals in both internal test set. Besides, we also made a further analysis to explore the prediction abilities of 5hmC biomarkers classifier in different stages in ESCC subgroups and found that there were significant differences between Early ESCC (stage 0 and I) and HC individuals, and the differences between stage I and HC samples ( P value = 4.54×10 − 3), stage 0 and stage I, II, III-IV were statistically significant too All the outcomes showed that the 5hmC indeed played an important role in tumor progress. Regretfully, in distinguishing stage 0 from HC, we failed to work out a precise identification ( P value = 0.40, with accuracy of 33.3%) which needed further investigation. So far, we had verified that the 5hmC could definitely serve as a favorable and non-invasive tool in differential diagnosis of esophageal cancer However, with the consideration of the importance of early detection of esophageal cancer and the deficient detection efficiency between early-stage esophageal cancer (stage 0) and HC samples in our study, we aim to establish a novel diagnostic model which can achieve our goals for early detection to satisfy clinical real needs. After employing low-pass whole-genome sequencing(WGS) technology, we worked out an integrated diagnostic classifier consisting of genome-wide 5’ end motif, nucleosome footprint (NF), fragmentation profiles and 5hmC biomarkers, and eventually realized an excellent AUC value of 0.934 with a sensitivity of 82.4%, specificity of 88.2% and accuracy of 84.3% for ESCC patient classification in the test set, especially the scores in ESCC patients with stage 0 and 1 were significantly higher than that in HC subjects( P values = 1.10×10 − 2 and 3.31×10 − 5 , respectively). Briefly speaking, our findings of combination of low-pass WGS of cfDNA signatures and 5hmC biomarkers improved the classifier’s efficiency from 65–82% of sensitivity at the specificity of 88% on an overall level. Most importantly, for stage 0 patients who had low-disease burden, the combined classifier significantly improved the prediction accuracy from 33.3–80.0%. Though the test cohort population in this study was still limited, it was still a valuable attempt for using combined approaches to improve prediction accuracy in non-invasion early diagnosis of ESCC. In general, more and more studies showed that integrating multi-omics detection was a promising methodology for non-invasive early diagnosis of many types of cancer. 5hmC was an intermediate produced through oxidation of 5-methylcytosine (5mC) by the oxygenase ten-eleven translocation (TET) enzyme family during 5mC demethylation process[ 39 – 41 ]. Both 5mC and 5hmC were presumed to have an important role in gene expression and regulation, and their modification changes were both observed in a wide range of malignant tumors, including ESCC[ 42 – 46 ]. Therefore, from a diagnostic perspective, combination of 5mC and 5hmC signal detection, as well as WGS features of cfDNA in this study to construct classifiers to further improve the specificity and sensitivity for early diagnosis of different subtypes and stages of ESCC patients was worthy of attempts. Here, an integrated model was worked out to realize early detection for esophageal cancer in a non-invasive approach that could be conveniently applied to clinical practices. Of course, further investigations about the stability of this model, the discriminating capabilities for different subtypes of esophageal cancer, or the practical application values are needed to execute. We sincerely hope that this study could provide an innovative clinical diagnostic strategy and could be put into practical use in the future which ultimately bring our patients with positive benefits. Conclusions This study suggests that the blood-based 5hmC integrated with low-pass WGS model could improve the accurate diagnosis of early stage ESCC, particularly for diagnosis of very early stage ESCC. Methods Study participants and clinical features All the patients selected to experimental group were diagnosed with esophageal squamous cell carcinoma (ESCC) and were confirmed cytopathologically and histologically, also, the patients were restricted to which on initial treatments from 0-PIV diagnosed by using the esophagus and esophagogastric junction of the eighth edition of the AJCC/UICC cancer staging manuals[ 47 ]. Except for living habits (especially hot food preference) and family disease history, other hazard covariates like BMI, smoking and drinking were kept consistent to the greatest extent. For patients subjected to control group, participants were selected from health examine center of our hospital, and no esophageal squamous cell carcinoma and other relevant diseases were observed. The other standards were consistent with experimental group. We excluded patients who used to receive surgery, chemoradiotherapy or immunotherapy, and patients who suffered from other illnesses like leukemia, neurodegenerative disease or other tumoral disorders. In total, 150 participants consisting of 100 esophageal squamous cell carcinoma (ESCC) patients and 50 healthy controls (HC) were enrolled from Nanfang Hospital of Southern Medical University in March 2019 to December 2020 and named as a Southern ESCC cohort since the patients were mainly from the Southern of China. Moreover, we added additional 21 healthy control samples used in a previously published study[ 25 ] to our HC cohort. Blood Sample Preparation And Cfdna Extraction Peripheral blood specimens (10 mL/subject) were obtained from patients who were newly diagnosed and had never received any medication or radical treatment for disease, which was prior to any biopsy or surgical resection. 71 HC blood samples were collected at the time of visiting the clinic for routine physical examination. All peripheral blood samples were stored in cell-free tubes (Streck, USA) at 4°C for no more than 72 hours before being separated into plasma and stored at -80°C in the laboratory. The plasma cell free DNA (cfDNA) was isolated using the MagMAX Cell-Free DNA Isolation Kit (Thermo, USA) according to the manufacturer's protocol. The quality of purified DNA was quantified by Qubit® 4.0 Fluorometer (Life Technologies, USA), and the DNA fragment size composition was assayed by Fragment Analyzer (Agilent, USA). 5hmc Sequencing And Data Processing 5hmC library construction and sequencing 5hmC library construction was performed according to the method previously described[ 23 ]. Briefly, the purified cfDNA (5–20 ng) were end-repaired and A tailed (5X ER/A-Tailing Enzyme Mix, Enzymatics, USA) then ligated with T-adaptors on both ends (WGS Ligase, Enzymatics, USA), to result in pre-library. Subsequently, ligated DNA was incubated in a 25 µl solution containing 50 mM HEPES buffer (pH = 8.0), 25 mM MgCl2, 60 µM UDP-6-N3-Glc (Active Motif, USA) and 12.5 U βGT (Thermo, USA) for 2 h at 37°C. Then, 2.5 µl DBCO-PEG4-biotin (Click Chemistry Tools, USA) was added to the reaction mixture and incubated for 2 h at 37°C. The purified DNA was incubated with 0.5 µl M270 Streptavidin beads (Life Technologies, USA) pre-blocked with salmon sperm DNA in buffer 1 (5 mM Tris pH 7.5, 0.5 mM EDTA, 1 M NaCl and 0.2% Tween 20) for 30 min. Afterwards, DNA fragments containing 5hmC features were subjected to PCR amplification, followed by the purification of the PCR products using AMPure XP beads according to the manufacturer's instructions. Finally, sequenced using Illumina NovaSeq 6000 platform. Mapping And Sequencing Quality Control The raw sequencing reads were removed adaptor and end sequence by trim_galore software (https: // www.bioinformatics . babraham. ac.uk/projects/trim_galore/)[ 48 ]. Acquired clean data were aligned to the human reference genome (hg19/GRCh37) by Bowtie2 v2.2.5 ( http://bowtiebio.sourceforge.net/bowtie2/index.shtml )[ 49 ]. Picard Tools ( http://broadinstitute.github.io/picard/ ) and SAMtools ( http://samtools.sourceforge.net/ )[ 50 ] were used to process and filter PCR duplicates for mapped BAM files. Reads with duplicate ratio less than 65% and enrichment efficiency over 95-fold passed the quality control and were used for further analysis. 5hmc Peak Identification Model-based analysis of ChIP-seq[ 51 ] was used to identify the 5hmC-enriched regions in each sample (the q value cut-off to call significant regions is 0.01; model fold = [ 5 , 50 ]). The peaks with high enrichment and significance (q 8) in all samples were considered as highly reliable 5hmC-enriched peaks. The 5hmC enrichment level was expressed as fragments per kilobase of 5hmC-DNA per million fragments mapped (FPKM). The genomic annotation of 5hmC peak regions was performed using annotatr[ 52 ] and the genome-wide distribution of 5hmC was visualized using the Integrated Genomics Viewer[ 53 , 54 ]. The metagene profile was generated using ngsplot[ 55 ]. Differential 5hmc Peak Regions Detection The differential 5hmC peak regions between healthy control group and ESCC group were identified using DESeq2 package[ 56 ] with P value < 0.01. De novo motif analysis around differential 5hmC peaks was performed using HOMER software (version 4.9). Functional gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses were performed using an online tool of metascape ( http://metascape.org/ )[ 57 ]. 5hmc Biomarkers Identification And Evaluating Performance 5hmC candidate biomarkers for cancer prediction models were identified firstly based on Wilcoxon rank-sum test ( P values < 0.001) between ESCC and HC groups and then processed to reduce the number of 5hmC biomarkers based on Recursive Feature Elimination - Cross Validation (RFECV) approach in the training set (54 HC, 34 Early ESCC and 31 Mid-Ad ESCC). The remaining samples in each group were used as the internal test set (including 17 HC, 16 Early ESCC and 19 Mid-Ad ESCC). In addition, we downloaded 150 esophageal cancer plasma-5hmC data and 183 healthy control plasma-5hmC data from the article previously published (nominated as the Northern ESCC cohort)[ 15 ] as an external test set to validate our results. The selected 273 differential 5hmC biomarkers using for sample identification to be ESCC or HC were analyzed by the principal component analysis (PCA). A heatmap of R package ( https://cran.r-project.org/web/packages/pheatmap/ index.html) was used to visualize hierarchical clustering and the distance in a heatmap figure[ 58 ]. 5hmC biomarkers were further processed for classification model construction based on a supervised two-class Support Vector Machine (SVM) method. To optimize the classification model and ensure the significance of each potential 5hmC biomarkers, the GridSearchCV package in Python in conjunction with cross-validation were performed to obtain the optimal parameters for SVM based on the Gaussian kernel (kernel='rbf', gamma = 0.001, the penalty parameter C = 4). The defined 5hmC-DNA regions and their corresponding genes were finally applied to classify the test set samples. Low-pass Whole Genome Sequencing And Data Processing WGS library construction and sequencing 1–10 ng cfDNA were end-repaired and A-tailed (Berry, China) then ligated with T-adaptors (Berry, China), to result in pre-library. The pre-libraries were purified by Clean NGS beads (VdoBiotech, China), followed by being quantified by the KAPA Library Quantification Kit (Kapa Biosystems, USA). cfDNA fragment size was confirmed using Bioanalyzer (Agilent, USA). Sequencing libraries were pooled at equal amount and then sequenced on an Illumina CN500 platform (Illumina, San Diego, USA) with an average coverage of 2x. Mapping And Sequencing Quality Control The raw sequencing reads were removed adaptor and end sequence by fastp software ( https://github.com/OpenGene/fastp ). Acquired clean data were aligned to human reference genome (hg19/GRCh37) using bwa-mem ( https://github.com/lh3/bwa ). Duplicate reads were marked by sambamba ( https://github.com/biod/sambamba/ ). Sequencing data were further processed by SAMtools ( http://samtools.sourceforge.net/ )[ 50 ] to get rid of marked duplicates, unmapped reads and low quality reads. Reads with duplicate rate less than 15% and mapping rate more than 95% passed the quality control and were used for further analysis. Wgs-based Biomarkers Identification And Integrated Model Construction To select more effective biomarkers for distinguishing ESCC samples from healthy controls, all samples were randomly separated into two subsets: the training set consisted of 54 HC, 30 Early ESCC and 29 Mid-Ad ESCC, and the test set consisted of the remaining samples (including 17 HC, 15 Early ESCC and 19 Mid-Ad ESCC). The Wilcoxon rank-sum test was used to compare biomarker features between ESCC and HC groups. Least Absolute Shrinkage and Selection Operator (LASSO) methods were applied to further reduce the number of biomarkers in the training set. The detailed selecting process was performed as follows. 5’ end Motif : 256 different types of 4mer 5’ end motif were identified and calculated their percentages (using pysam ( https://pysam.readthedocs.io/en/latest/ )) without considering chromosome Y and unidentifiable bases. Following motif types were filtered out: 1) P ≥ 0.05 in Wilcoxon rank-sum test between ESCC and HC groups; 2) weight of 0 via LASSO. Eventually, 120 motif types were left for further analysis. Nucleosome footprint (NF) We obtained all transcripts of coding genes, microRNAs and long non-coding RNAs (LncRNAs) and calculated the distance between transcripts. The transcripts with the distance more than 200 bp were retained. If the distance between two transcripts less than 200 bp, the longer transcript was retained. A total of 57151 transcripts of 30588 genes were recruited for analysis. The promoter region and background region of transcripts were divided, and the reads number of different regions was counted with featureCounts[ 59 ]. NF score of each gene is calculated as NF Score=(background1 + background2)/2-Promotor Following genes were filtered out: 1) more than 10% of the total samples showing an NF score of 0; 2) P ≥ 0.001 in Wilcoxon rank-sum test between ESCC and HC groups; 3) weight of 0 via LASSO. Eventually, 170 genes were left for further analysis. Fragment : The whole genome except Y chromosome was divided into 1M sized bins, resulting in 3055 areas. Pysam ( https://pysam.readthedocs.io/en/latest/ ) was used to calculate the length of insertion fragment and ratio of short/long fragment in different regions. LASSO was then used to filter out areas with a weight of 0, and finally 10 areas were retained. Thereafter, the Support Vector Machine (SVM) method was implemented for individual genomic feature-based model construction. 10-fold cross-validation method was employed to optimize the combination of the parameters in the training set, and cut-off value was set at the point with the best diagnostic accuracy. To obtain the best diagnostic model, logistic regression model was generated using the predictive score of the four individual models as input features to integrate the outcome of each model based on the training dataset. The logistic Score was calculated as follows. Logistic Score = exp(Z) / (1 + exp(Z)), where Z= -2.57+(3.35×NF)+༈0.05×Fragment༉+༈0.75×Motif༉+༈1.74×5hmC༉ Receiver operating characteristic (ROC) curves[ 60 ] were generated to evaluate the performance of a prediction algorithm, using the pROC[ 61 ] library in the R package. Sensitivity and specificity were estimated at the score cut-off that maximizes the sum of sensitivity and specificity using the ROCR library in the R package. Statistics The statistical methods used were stated in the above methods. The R code related to classifier detection and modelling is available upon request. Declarations Ethics approval and consent to participate The study was approved by the ethics committee of the Nanfang Hospital, Southern Medical University, Guangzhou, China (reference: NFEC-2019-014) and was also registered with ClinicalTrials.gov (reference: NCT03922230). Besides, this study was conducted by the approval of the Institutional Review Board of Nanfang Hospital of Southern Medical University, and the written informed consents were obtained from all participants according to the institutional guidelines. Consent for publication All the authors agreed to submit and publish the manuscript to Journal of Hematology & Oncology. Data availability statement All of the raw and processed data used in this study have been uploaded to the Genome Sequence Archive depository (https://ngdc.cncb.ac.cn/gsa)[ 62 ] with the Accession Number (HRA001476). Conflict of interest: SW, HL, FS, SW and XZ are employees of Berry Oncology Corporation.Other authors had no declaration of conflicts of interests. Funding . Science and Technology Planning Project of Guangdong Province(2017B020226005). Author contributions KC, SW and DLdesigned the research. DL,XW, JZ and XD recruited the examinees. DL, XW and SF got consents and collected raw data. DL, XY, HL and SWwrote the manuscript. SW, XZ,FS conducted bioinformatics analysisanddata visualization. All authors reviewed and approved the manuscript. Acknowledgments Thanks for the help of Professor Side Liu, Dr. Jianqun Cai and Dr. Jing Wang from Department of Gastroenterology, Ms. Li Zhen from Department of General Surgery, Ms. Yu Guo from Department of Huiqiao Building, and Ms. Mei Li from Department of Thoracic Surgery, Nanfang Hospital, Southern Medical University, Guangzhou, China. This study was funded by the Science and Technology Planning Project of Guangdong Province (2017B020226005). References Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F: Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries . CA: a cancer journal for clinicians 2021, 71 (3):209-249. Smyth EC, Lagergren J, Fitzgerald RC, Lordick F, Shah MA, Lagergren P, Cunningham D: Oesophageal cancer . Nature reviews Disease primers 2017, 3 :17048. Allemani C, Matsuda T, Di Carlo V, Harewood R, Matz M, Nikšić M, Bonaventure A, Valkov M, Johnson CJ, Estève J et al : Global surveillance of trends in cancer survival 2000-14 (CONCORD-3): analysis of individual records for 37 513 025 patients diagnosed with one of 18 cancers from 322 population-based registries in 71 countries . Lancet 2018, 391 (10125):1023-1075. Wani S, Yadlapati R, Singh S, Sawas T, Katzka DA, Hall M, Bergman J, Canto MI, Chak A, Corley DA et al : Post-Endoscopy Esophageal Neoplasia in Barrett’s Esophagus: Consensus Statements from an International Expert Panel . Gastroenterology 2021. di Pietro M, Canto MI, Fitzgerald RC: Endoscopic Management of Early Adenocarcinoma and Squamous Cell Carcinoma of the Esophagus: Screening, Diagnosis, and Therapy . Gastroenterology 2018, 154 (2):421-436. Wan JCM, Massie C, Garcia-Corbacho J, Mouliere F, Brenton JD, Caldas C, Pacey S, Baird R, Rosenfeld N: Liquid biopsies come of age: towards implementation of circulating tumour DNA . Nature reviews Cancer 2017, 17 (4):223-238. Thierry AR, El Messaoudi S, Gahan PB, Anker P, Stroun M: Origins, structures, and functions of circulating DNA in oncology . Cancer metastasis reviews 2016, 35 (3):347-376. Cristiano S, Leal A, Phallen J, Fiksel J, Adleff V, Bruhm DC, Jensen SØ, Medina JE, Hruban C, White JR et al : Genome-wide cell-free DNA fragmentation in patients with cancer . Nature 2019, 570 (7761):385-389. Luo H, Wei W, Ye Z, Zheng J, Xu RH: Liquid Biopsy of Methylation Biomarkers in Cell-Free DNA . Trends in molecular medicine 2021, 27 (5):482-500. Jones PA: Functions of DNA methylation: islands, start sites, gene bodies and beyond . Nature reviews Genetics 2012, 13 (7):484-492. Baylin SB, Jones PA: A decade of exploring the cancer epigenome - biological and translational implications . Nature reviews Cancer 2011, 11 (10):726-734. Dor Y, Cedar H: Principles of DNA methylation and their implications for biology and medicine . The Lancet 2018, 392 (10149):777-786. Branco MR, Ficz G, Reik W: Uncovering the role of 5-hydroxymethylcytosine in the epigenome . Nature reviews Genetics 2011, 13 (1):7-13. Guler GD, Ning Y, Ku CJ, Phillips T, McCarthy E, Ellison CK, Bergamaschi A, Collin F, Lloyd P, Scott A et al : Detection of early stage pancreatic cancer using 5-hydroxymethylcytosine signatures in circulating cell free DNA . Nature communications 2020, 11 (1):5270. Zhang J, Han X, Gao C, Xing Y, Qi Z, Liu R, Wang Y, Zhang X, Yang YG, Li X et al : 5-Hydroxymethylome in Circulating Cell-free DNA as A Potential Biomarker for Non-small-cell Lung Cancer . Genomics, proteomics & bioinformatics 2018, 16 (3):187-199. Gilat N, Tabachnik T, Shwartz A, Shahal T, Torchinsky D, Michaeli Y, Nifker G, Zirkin S, Ebenstein Y: Single-molecule quantification of 5-hydroxymethylcytosine for diagnosis of blood and colon cancers . Clinical epigenetics 2017, 9 :70. Li W, Zhang X, Lu X, You L, Song Y, Luo Z, Zhang J, Nie J, Zheng W, Xu D et al : 5-Hydroxymethylcytosine signatures in circulating cell-free DNA as diagnostic biomarkers for human cancers . Cell research 2017, 27 (10):1243-1257. Song CX, Yin S, Ma L, Wheeler A, Chen Y, Zhang Y, Liu B, Xiong J, Zhang W, Hu J et al : 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages . Cell research 2017, 27 (10):1231-1242. Cai J, Chen L, Zhang Z, Zhang X, Lu X, Liu W, Shi G, Ge Y, Gao P, Yang Y et al : Genome-wide mapping of 5-hydroxymethylcytosines in circulating cell-free DNA as a non-invasive approach for early detection of hepatocellular carcinoma . Gut 2019, 68 (12):2195-2205. Zviran A, Schulman RC, Shah M, Hill STK, Deochand S, Khamnei CC, Maloney D, Patel K, Liao W, Widman AJ et al : Genome-wide cell-free DNA mutational integration enables ultra-sensitive cancer monitoring . Nat Med 2020, 26 (7):1114-1124. Chabon JJ, Hamilton EG, Kurtz DM, Esfahani MS, Moding EJ, Stehr H, Schroers-Martin J, Nabet BY, Chen B, Chaudhuri AA et al : Integrating genomic features for non-invasive early lung cancer detection . Nature 2020, 580 (7802):245-251. Mathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, Bruhm DC, Niknafs N, Ferreira L, Adleff V et al : Detection and characterization of lung cancer using cell-free DNA fragmentomes . Nature communications 2021, 12 (1):5060. Chen L, Abou-Alfa GK, Zheng B, Liu JF, Bai J, Du LT, Qian YS, Fan R, Liu XL, Wu L et al : Genome-scale profiling of circulating cell-free DNA signatures for early detection of hepatocellular carcinoma in cirrhotic patients . Cell research 2021, 31 (5):589-592. Szymanski JJ, Sundby RT, Jones PA, Srihari D, Earland N, Harris PK, Feng W, Qaium F, Lei H, Roberts D et al : Cell-free DNA ultra-low-pass whole genome sequencing to distinguish malignant peripheral nerve sheath tumor (MPNST) from its benign precursor lesion: A cross-sectional study . PLoS Med 2021, 18 (8):e1003734. Xiao Z, Wu W, Wu C, Li M, Sun F, Zheng L, Liu G, Li X, Yun Z, Tang J et al : 5-Hydroxymethylcytosine signature in circulating cell-free DNA as a potential diagnostic factor for early-stage colorectal cancer and precancerous adenoma . Molecular oncology 2021, 15 (1):138-150. Tian X, Sun B, Chen C, Gao C, Zhang J, Lu X, Wang L, Li X, Xing Y, Liu R et al : Circulating tumor DNA 5-hydroxymethylcytosine as a novel diagnostic biomarker for esophageal cancer . Cell research 2018, 28 (5):597-600. Oikawa T: ETS transcription factors: possible targets for cancer therapy . Cancer science 2004, 95 (8):626-633. Romano O, Miccio A: GATA factor transcriptional activity: Insights from genome-wide binding profiles . IUBMB Life 2020, 72 (1):10-26. Chen D, Wang K, Li X, Jiang M, Ni L, Xu B, Chu Y, Wang W, Wang H, Kang H et al : FOXK1 plays an oncogenic role in the development of esophageal cancer . Biochem Biophys Res Commun 2017, 494 (1-2):88-94. Serpas L, Chan RWY, Jiang P, Ni M, Sun K, Rashidfarrokhi A, Soni C, Sisirak V, Lee W-S, Cheng SH et al : deletion causes aberrations in length and end-motif frequencies in plasma DNA . Proc Natl Acad Sci U S A 2019, 116 (2):641-649. Zhao Y, Wang J, Liang F, Liu Y, Wang Q, Zhang H, Jiang M, Zhang Z, Zhao W, Bao Y et al : NucMap: a database of genome-wide nucleosome positioning map across species . Nucleic Acids Res 2019, 47 (D1):D163-D169. Wei GH, Badis G, Berger MF, Kivioja T, Palin K, Enge M, Bonke M, Jolma A, Varjosalo M, Gehrke AR et al : Genome-wide analysis of ETS-family DNA-binding in vitro and in vivo . The EMBO journal 2010, 29 (13):2147-2160. Dittmer J: The role of the transcription factor Ets1 in carcinoma . Seminars in cancer biology 2015, 35 :20-38. Saeki H, Kuwano H, Kawaguchi H, Ohno S, Sugimachi K: Expression of ets-1 transcription factor is correlated with penetrating tumor progression in patients with squamous cell carcinoma of the esophagus . Cancer 2000, 89 (8):1670-1676. Saeki H, Oda S, Kawaguchi H, Ohno S, Kuwano H, Maehara Y, Sugimachi K: Concurrent overexpression of Ets-1 and c-Met correlates with a phenotype of high cellular motility in human esophageal cancer . Int J Cancer 2002, 98 (1). Baltrunaite K, Craig MP, Palencia Desai S, Chaturvedi P, Pandey RN, Hegde RS, Sumanas S: ETS transcription factors Etv2 and Fli1b are required for tumor angiogenesis . Angiogenesis 2017, 20 (3):307-323. Adamo P, Ladomery MR: The oncogene ERG: a key factor in prostate cancer . Oncogene 2016, 35 (4):403-414. Qiao G, Zhuang W, Dong B, Li C, Xu J, Wang G, Xie L, Zhou Z, Tian D, Chen G et al : Discovery and validation of methylation signatures in circulating cell-free DNA for early detection of esophageal cancer: a case-control study . BMC medicine 2021, 19 (1):243. Gilat N, Tabachnik T, Shwartz A, Shahal T, Torchinsky D, Michaeli Y, Nifker G, Zirkin S, Ebenstein Y: Single-molecule quantification of 5-hydroxymethylcytosine for diagnosis of blood and colon cancers . Clinical epigenetics 2017, 9 :70. Song C-X, Yin S, Ma L, Wheeler A, Chen Y, Zhang Y, Liu B, Xiong J, Zhang W, Hu J et al : 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages . Cell research 2017, 27 (10):1231-1242. Li W, Zhang X, Lu X, You L, Song Y, Luo Z, Zhang J, Nie J, Zheng W, Xu D et al : 5-Hydroxymethylcytosine signatures in circulating cell-free DNA as diagnostic biomarkers for human cancers . Cell research 2017, 27 (10):1243-1257. Shi X, Yu Y, Luo M, Zhang Z, Shi S, Feng X, Chen Z, He J: Loss of 5-Hydroxymethylcytosine Is an Independent Unfavorable Prognostic Factor for Esophageal Squamous Cell Carcinoma . PloS one 2016, 11 (4):e0153100. Murata A, Baba Y, Ishimoto T, Miyake K, Kosumi K, Harada K, Kurashige J, Iwagami S, Sakamoto Y, Miyamoto Y et al : TET family proteins and 5-hydroxymethylcytosine in esophageal squamous cell carcinoma . Oncotarget 2015, 6 (27):23372-23382. Li D, Zhang L, Liu Y, Sun H, Onwuka JU, Zhao Z, Tian W, Xu J, Zhao Y, Xu H: Specific DNA methylation markers in the diagnosis and prognosis of esophageal cancer . Aging (Albany NY) 2019, 11 (23):11640-11658. Su J, Wu G, Ye Y, Zhang J, Zeng L, Huang X, Zheng Y, Bai R, Zhuang L, Li M et al : NSUN2-mediated RNA 5-methylcytosine promotes esophageal squamous cell carcinoma progression via LIN28B-dependent GRB2 mRNA stabilization . Oncogene 2021, 40 (39):5814-5828. Kit OI, Vodolazhskiy DI, Kolesnikov EN, Timoshkina NN: [Epigenetic markers of esophageal cancer: DNA methylation] . Biomed Khim 2016, 62 (5):520-526. Rice TW, Ishwaran H, Ferguson MK, Blackstone EH, Goldstraw P: Cancer of the Esophagus and Esophagogastric Junction: An Eighth Edition Staging Primer . J Thorac Oncol 2017, 12 (1):36-42. Lindgreen S: AdapterRemoval: easy cleaning of next-generation sequencing reads . BMC Res Notes 2012, 5 :337. Langmead B, Salzberg SL: Fast gapped-read alignment with Bowtie 2 . Nat Methods 2012, 9 (4):357-359. Etherington GJ, Ramirez-Gonzalez RH, MacLean D: bio-samtools 2: a package for analysis and visualization of sequence and alignment data with SAMtools in Ruby . Bioinformatics 2015, 31 (15):2565-2567. Grytten I, Rand KD, Nederbragt AJ, Storvik GO, Glad IK, Sandve GK: Graph Peak Caller: Calling ChIP-seq peaks on graph-based reference genomes . PLoS Comput Biol 2019, 15 (2):e1006731. Cavalcante RG, Sartor MA: annotatr: genomic regions in context . Bioinformatics 2017, 33 (15):2381-2383. Thorvaldsdóttir H, Robinson JT, Mesirov JP: Integrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration . Brief Bioinform 2013, 14 (2):178-192. Robinson JT, Thorvaldsdóttir H, Winckler W, Guttman M, Lander ES, Getz G, Mesirov JP: Integrative genomics viewer . Nat Biotechnol 2011, 29 (1):24-26. Shen L, Shao N, Liu X, Nestler E: ngs.plot: Quick mining and visualization of next-generation sequencing data by integrating genomic databases . BMC Genomics 2014, 15 :284. Love MI, Huber W, Anders S: Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2 . Genome Biol 2014, 15 (12):550. Tripathi S, Pohl MO, Zhou Y, Rodriguez-Frandsen A, Wang G, Stein DA, Moulton HM, DeJesus P, Che J, Mulder LCF et al : Meta- and Orthogonal Integration of Influenza "OMICs" Data Defines a Role for UBR4 in Virus Budding . Cell Host Microbe 2015, 18 (6):723-735. Galili T, O'Callaghan A, Sidi J, Sievert C: heatmaply: an R package for creating interactive cluster heatmaps for online publishing . Bioinformatics 2018, 34 (9):1600-1602. Liao Y, Smyth GK, Shi W: featureCounts: an efficient general purpose program for assigning sequence reads to genomic features . Bioinformatics 2014, 30 (7):923-930. Sing T, Sander O, Beerenwinkel N, Lengauer T: ROCR: visualizing classifier performance in R . Bioinformatics 2005, 21 (20):3940-3941. Robin X, Turck N, Hainard A, Tiberti N, Lisacek F, Sanchez J-C, Müller M: pROC: an open-source package for R and S+ to analyze and compare ROC curves . BMC Bioinformatics 2011, 12 :77. Wang Y, Song F, Zhu J, Zhang S, Yang Y, Chen T, Tang B, Dong L, Ding N, Zhang Q et al : GSA: Genome Sequence Archive . Genomics, proteomics & bioinformatics 2017, 15 (1):14-18. Additional Declarations Competing interest reported. SW, HL, FS, SW and XZ are employees of Berry Oncology Corporation. Other authors had no declarations on conflicts of interests. All the authors agreed to submit and publish the article in Journal of Hematology & Oncology. Supplementary Files SupplementaryLegendsandfiguresfinal.pdf TableS1Clinicopathologicalcharacteristics.xlsx TableS2Thetop10significantlyenrichment5hmCmotifs.xlsx TableS3Listofannotated5hmCmarkergenesinmodel.xlsx TableS4ListofWGSfeaturesinmodel.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1375061","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":87001020,"identity":"4446a57e-d9c2-4087-aa5e-0df88d22ad3b","order_by":0,"name":"Di Lu","email":"","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Di","middleName":"","lastName":"Lu","suffix":""},{"id":87001021,"identity":"43e0a273-7c5e-4328-8f8b-84426053f132","order_by":1,"name":"Xuanzhen Wu","email":"","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xuanzhen","middleName":"","lastName":"Wu","suffix":""},{"id":87001022,"identity":"ad3c88bc-a4a3-4b20-b5e4-708f835fb788","order_by":2,"name":"Shuangxiu Wu","email":"","orcid":"","institution":"Berry Oncology Corporation","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shuangxiu","middleName":"","lastName":"Wu","suffix":""},{"id":87001023,"identity":"e5c4739e-0c80-4557-ad6e-4eee05eaee70","order_by":3,"name":"Hui Li","email":"","orcid":"","institution":"Berry Oncology Corporation","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hui","middleName":"","lastName":"Li","suffix":""},{"id":87001024,"identity":"3f9d5750-7015-4bea-a330-f0d03327c7b3","order_by":4,"name":"Xuebin Yan","email":"","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xuebin","middleName":"","lastName":"Yan","suffix":""},{"id":87001025,"identity":"330674f3-19b8-4b56-946d-2ad01a7899a6","order_by":5,"name":"Jianxue Zhai","email":"","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jianxue","middleName":"","lastName":"Zhai","suffix":""},{"id":87001026,"identity":"bce53628-7a14-42bb-8518-b9799e8bb1ee","order_by":6,"name":"Xiaoying Dong","email":"","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xiaoying","middleName":"","lastName":"Dong","suffix":""},{"id":87001027,"identity":"63866840-64d8-4fc3-8fa7-0cfc80d32c10","order_by":7,"name":"Siyang Feng","email":"","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Siyang","middleName":"","lastName":"Feng","suffix":""},{"id":87001028,"identity":"a50ba69d-ab80-422b-b213-0d3cbd5f9a53","order_by":8,"name":"Fuming Sun","email":"","orcid":"","institution":"Berry Oncology Corporation","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Fuming","middleName":"","lastName":"Sun","suffix":""},{"id":87001029,"identity":"32f5eac5-84dc-41e5-ac23-7e9692fd2847","order_by":9,"name":"Shaobo Wang","email":"","orcid":"","institution":"Berry Oncology Corporation","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shaobo","middleName":"","lastName":"Wang","suffix":""},{"id":87001030,"identity":"614c990c-429a-490e-b178-b958b4256ecf","order_by":10,"name":"Xueying Zhang","email":"","orcid":"","institution":"Berry Oncology Corporation","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xueying","middleName":"","lastName":"Zhang","suffix":""},{"id":87001031,"identity":"52a2bec5-d194-433f-9e50-5300f69a2c32","order_by":11,"name":"Kaican Cai","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAo0lEQVRIiWNgGAWjYDCCA2BSQo6Nvf0AaVqM+XjOJJCkhSFxnoSDAXE6+I43P2D4UWOR3ibBkMDwo2IbYS2SZ44ZMPYck8htk248wNhz5jZhLQY3chiYGdiAWmQOJDAzthGt5Z9EOptEggEJWhjbJBKI1wL2S2+fhGEbMJAPEuUXSIh9q5OXb28/+OBHBRFagID9B4x1gCj1o2AUjIJRMAoIAwAPqjhhf/c4nQAAAABJRU5ErkJggg==","orcid":"","institution":"Nanfang Hospital, Southern Medical University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Kaican","middleName":"","lastName":"Cai","suffix":""}],"badges":[],"createdAt":"2022-02-19 04:44:07","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1375061/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1375061/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":18791137,"identity":"cf2fdcbb-c6b1-42c7-baf2-1b73bc9afb83","added_by":"auto","created_at":"2022-03-02 17:49:36","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":150722,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSketch map of study design and research pipeline for early detection of ESCC. \u003c/strong\u003e\u003c/p\u003e\u003cp\u003e5hmC-based diagnostic model and low-pass WGS-based diagnostic model were developed for identifying ctDNA from plasma cfDNA using machine learning approach.Totally 171 subjects were involved as a Southern ESCC cohort and blood samples were collected to perform 5hmC-seqeuncing and low-pass WGS, respectively. Two-thirds subjects were randomly selected as a training set and the left one-third subjects were used as an independently internal Southern-ESCC test set to evaluate the model performance. A downloaded ESCC-5hmC dataset was used as an independently external Northern-ESCC test set. Detail research pipeline was illustrated in supplementary Figure S2.\u003c/p\u003e\u003cp\u003ecfDNA, cell-free tumor DNA. cfDNA, cell-free DNA; HC, healthy controls individuals; ESCC, esophageal squamous cell carcinoma; Mid-Ad, middle-advanced; 5hmC, 5-hydroxymethylcytosines; WGS, whole genome sequencing.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/9091630d121d299a08e75125.png"},{"id":18791001,"identity":"f1d92dc3-c6de-485a-992a-ba48cd82aa73","added_by":"auto","created_at":"2022-03-02 17:46:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":349649,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenome-wide distribution of 5hmC signals in plasma cfDNA of ESCC and HC individuals. \u003c/strong\u003e\u003c/p\u003e\u003cp\u003e(A) Comparison of density distribution of 5hmC peaks number in plasma samples from 71 HC and 100 patients with ESCC. \u003c/p\u003e\u003cp\u003e(B) Comparison of the total number of 5hmC peaks in HC and ESCC patients with stage 0-I, II, III-IV. Each dot depicts an individual cfDNA sample. \u003cem\u003eP\u003c/em\u003e value shows statistical significance by Mann-Kendall Test. \u003c/p\u003e\u003cp\u003e(C) Metagene profiles of mean values of 5hmC read counts on the regions from TSS to TES with the flanking 3000-bp in HC and ESCC samples. \u003c/p\u003e\u003cp\u003e(D) Distribution of differential 5hmC peaks in genomic elements in ESCC samples versus HC samples. \u003c/p\u003e\u003cp\u003e(E) Top enriched known transcription factor binding motifs detected in differential 5hmC peaks (left: 5hmC up-regulation; right: 5hmC down-regulation). Motif information was obtained from the Homer motif database. The value in parenthesis represents the percentage of target sequences enriched with the binding motif of the indicated transcription factor. \u003c/p\u003e\u003cp\u003e(F) Venn plot of the top 10-significantly enrichment 5hmC motifs in our Southern ESCC group, the Northern ESCC group and the CRC group (left: 5hmC up-regulation; right: 5hmC down-regulation). \u003c/p\u003e\u003cp\u003eHC, healthy controls; ESCC, esophageal squamous cell carcinoma; CRC, colorectal cancer; TSS, transcription\u0026nbsp;start\u0026nbsp;sites; TES, transcription end site; 5hmC, 5-hydroxymethylcytosines.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/55ac697b0341f56fb40f1ea2.png"},{"id":18791003,"identity":"a63f7cf6-7120-4f05-8fb1-5bed1010405e","added_by":"auto","created_at":"2022-03-02 17:46:36","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":630964,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDevelopment, validation and performance of 5hmC diagnostic model.\u003c/strong\u003e\u003c/p\u003e\u003cp\u003e(A) Unsupervised hierarchical clustering of 71 HC and 100 ESCC cfDNA samples based on top 273 5hmC marker genes. \u003c/p\u003e\u003cp\u003e(B) PCA plot of 71 HC and 100 ESCC cfDNA samples based on top 273 5hmC marker genes. \u003c/p\u003e\u003cp\u003e(C) GO enrichment (left) and KEGG pathway enrichment (right) analysis of 273 biomarkers of the 5hmC classifier.\u003c/p\u003e\u003cp\u003e(D) The normalized 5hmC values of \u003cem\u003eFOXK1\u003c/em\u003e in HC and ESCC samples. (E-F) ROC curves and associated AUC values in the internal test set (E) and the external test set (F). \u003c/p\u003e\u003cp\u003e(G) Predictive probability scores based on 5hmC classifier for different clinical stages of internal test set samples. \u003c/p\u003e\u003cp\u003eHC, healthy controls; ESCC, esophageal squamous cell carcinoma; \u003cem\u003eFOXK1\u003c/em\u003e, forkhead box K1; ROC, receiver operating characteristic; AUC, area under curve; 5hmC, 5-hydroxymethylcytosines.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/370401fd819b01269240de15.png"},{"id":18791004,"identity":"ded1540a-802d-405d-93d9-d56ce0974100","added_by":"auto","created_at":"2022-03-02 17:46:36","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":767398,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDevelopment, validation and performance of the integrated diagnostic model.\u003c/strong\u003e\u003c/p\u003e\u003cp\u003e(A) Heatmap analysis of differential motifs (\u003cem\u003ep\u003c/em\u003e value \u0026lt; 0.001) between ESCC and HC samples.\u003c/p\u003e\u003cp\u003e(B) Heatmap analysis of genes with differential reads coverage between gene promoter and background regions (\u003cem\u003ep\u003c/em\u003e values \u0026lt; 0.001)in ESCC and HC samples.\u003c/p\u003e\u003cp\u003e(C) Frequencies comparison of different fragment sizes between ESCC and HC samples.\u003c/p\u003e\u003cp\u003e(D) ROC curves and associated AUC values in the test set.\u003c/p\u003e\u003cp\u003e(E) Confusion matrices of integrated diagnostic model comparing the actual class with the predicted class for ESCC (n = 34) and HC (n = 17) samples in the test set.\u003c/p\u003e\u003cp\u003e(F) Predictive probability scores based on integrated diagnostic model for different clinical stages of the test set samples.\u003c/p\u003e\u003cp\u003e(G) Comparison of diagnostic performance between 5hmC model (blue) and integrated model (red) on different clinical stages of the test set ESCC samples (n = 35). The blue and red dotted line represent the thresholds of diagnostic positive of the 5hmC model and the integrated model, respectively. Positive ESCC detection is indicated by black dots, negative ESCC detection indicated by blank.\u003c/p\u003e\u003cp\u003eHC, healthy controls; ESCC, esophageal squamous cell carcinoma; ROC, receiver operating characteristic; AUC, area under curve; 5hmC, 5-hydroxymethylcytosines.\u003c/p\u003e\u003cp\u003e\u003cbr\u003e\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/ea1a90b1c50bafee9c52e144.png"},{"id":19052114,"identity":"f9bac027-c96c-426e-baca-97298c93cca9","added_by":"auto","created_at":"2022-03-09 20:29:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3823798,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/a27ab56d-7a65-4313-9554-e2fc23d9b9c3.pdf"},{"id":18791005,"identity":"c0f3cc72-341e-4c56-8720-9d0e6877b6a2","added_by":"auto","created_at":"2022-03-02 17:46:36","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":432350,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryLegendsandfiguresfinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/8be00b7eb2e9957fa90c3b45.pdf"},{"id":18791006,"identity":"d5e6de95-4212-41b2-82bc-9b71773c0f0f","added_by":"auto","created_at":"2022-03-02 17:46:36","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":24394,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1Clinicopathologicalcharacteristics.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/0e56925947f0b86553329b97.xlsx"},{"id":18791011,"identity":"5e7b389f-c625-4535-b3b0-de78366be32a","added_by":"auto","created_at":"2022-03-02 17:46:37","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":12042,"visible":true,"origin":"","legend":"","description":"","filename":"TableS2Thetop10significantlyenrichment5hmCmotifs.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/51855035318d9246b0f3d42e.xlsx"},{"id":18791007,"identity":"b17424e4-5018-4316-89ed-269b6d5f2058","added_by":"auto","created_at":"2022-03-02 17:46:37","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":39479,"visible":true,"origin":"","legend":"","description":"","filename":"TableS3Listofannotated5hmCmarkergenesinmodel.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/d0fc5b50cdb3025a0ac7c808.xlsx"},{"id":18791008,"identity":"2167ed69-17fb-4ac6-aa14-7567d5163c82","added_by":"auto","created_at":"2022-03-02 17:46:37","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":25877,"visible":true,"origin":"","legend":"","description":"","filename":"TableS4ListofWGSfeaturesinmodel.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-1375061/v1/823776e41ff23ed4c296725d.xlsx"}],"financialInterests":"Competing interest reported. SW, HL, FS, SW and XZ are employees of Berry Oncology Corporation. Other authors had no declarations on conflicts of interests. All the authors agreed to submit and publish the article in Journal of Hematology \u0026 Oncology.","formattedTitle":"\u003cp\u003ePlasma Cell-free DNA 5-Hydroxymethylcytosine and Whole-Genome Sequencing Signatures for Early Detection of Esophageal Cancer\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eEsophageal cancer is a global problem threatening people\u0026rsquo;s health and declining people\u0026rsquo;s life expectancy which ranks the seventh in terms of incidence and the sixth in mortality worldwide in 2020 with two main subtypes: esophageal squamous cell carcinoma (ESCC) and esophageal adenocarcinoma (EAC)[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].There are distinct geographic variations between the two main subgroups of esophageal cancer. The incidence of ESCC is apparently higher in China, Central Asia, East and South of Africa while most of the EAC cases occur in patients living in high income countries like North America and Europe[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].Most of the patients were already suffered from ESCC with an advanced stage when they sought medical advices as early esophageal cancer was rarely accompanied by distinct symptoms, thus, neither the surgical procedures nor adjuvant therapies could lead to satisfied outcomes[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. So, an urgent demand for a new method to help making early diagnosis of esophageal carcinoma is desired.\u003c/p\u003e \u003cp\u003eGenerally, tissue biopsy like endoscopy is the most widespread procedure used to make detection of esophageal cancer and help formulating stages for further therapeutic schedule. However, there remains obvious shortcomings and inadequacies like inferior diagnostic efficiency, uncomfortable checking process and expensive spending which may lead to poor adherence to subsequent treatments in return[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Compared to conventional tissue biopsy, liquid biopsy is an emerging non-invasive method to make early detection which can also make further molecular and genic analysis using these liquid specimens[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Thereinto, cell-free DNA (cfDNA) is deliberated from the cell existing in blood, lymph, bile, milk, urine, saliva, mucous suspension, spinal fluid and amniotic fluid originating from apoptosis, necrosis, phagocytosis, nocosis and active secretion[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], and was proved to be a approving biomarker to make screening, detection and monitoring of cancers[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. However, there remains some challenges in the application of cfDNA for early detection of cancers like hyposecretion, variability and polyphyly[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. The efficiency of DNA methylation as an epigenetic mark has been previously validated[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], furthermore, the association between DNA methylation and tumorigenesis has been verified too[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. As the conversion of cytosine, 5-hydroxymethylcytosine (5hmC) is oxidized from 5-methylcytosine(5mC) by the enzyme ten-eleven translocation 1(TET1) which is one of the three enzymes of TET(ten-eleven translocation) family and is recognized as a better marker to detect gene expression[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Although there are fewer 5hmC existing in tissues and more challenges to accomplish precise detection, it has been proved that 5hmC has a superior permissive effect on gene expression and exhibits more tissue specificity. Besides, plenty of studies have reported that 5hmC in cfDNA was used as a promising marker to detect cancers like early-stage pancreatic cancer, non-small-cell lung cancer (NSCLC), hepatocellular carcinoma (HCC), blood and colon cancer[\u003cspan additionalcitationids=\"CR15 CR16 CR17 CR18\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. In addition, a recent study showed ctDNA in plasma had distinct features, such as shorter fragment size, special motif and nucleosome footprint (NF) in comparison with cfDNA of healthy people[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], which could be applied to an advanced technology (low-pass whole-genome sequencing (WGS)) to get more accurate outcomes in identifying ctDNA from plasma. WGS of plasma cfDNA was a new empowering technology used for non-invasion low-disease-burden diagnosis. This technology provided ultra-sensitive detection of ctDNA from cfDNA of healthy individuals via genome-wide mutational integration[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Recent studies proved its good performance on early diagnosis of lung cancer[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], hepatocellular carcinoma (HCC)[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] and malignant peripheral nerve sheath tumor (MPNST)[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTherefore, employing 5hmC in plasma cfDNA as a biomarker to distinguish the differences between ESCC patients and healthy control people (HC) is worth trying, and we aimed to determine 5hmC signatures in cfDNA using machine learning methods as early diagnostic biomarkers for ECSS in a prospective cohort study. The low-pass WGS technology was also attempted as a complementary technique for identification of early ESCC from healthy people.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003eSamples composition and study design\u003c/h2\u003e\n\u003cp\u003eA total of 150 adult subjects were prospectively enrolled in this study from March 2019 to December 2020, including patients with Early ESCC (stages 0, IA and IB, n\u0026thinsp;=\u0026thinsp;50), and middle and advanced (Mid-Ad) ESCC (stages II, III and IVA, n\u0026thinsp;=\u0026thinsp;50) and healthy control individuals (HC, n\u0026thinsp;=\u0026thinsp;50) from Nanfang Hospital of Southern Medical University, as a Southern ESCC cohort since the patients were mainly from the South of China. Detailed clinical information regarding gender, age, BMI, living habits, differentiation of cancer, TNM stages and surgery selections were illustrated in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e. And after conducting a brief statistical analysis using the data of the 150 patients collected from our hospital, we observed that there were no apparent differences among most of the variates like gender, age, BMI, habits and the incidences of hypertension and diabetes (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.05), except for family history (\u003cem\u003eP\u0026thinsp;\u0026lt;\u003c/em\u003e\u0026thinsp;0.05), which suggested that family history might play an important role in tumorigenesis.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eSummary of demographic and clinicopathological characteristics of all the participants in this study\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eCharacteristics\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eHC\u003c/p\u003e\n\u003cp\u003e(n\u0026thinsp;=\u0026thinsp;50)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eEarly ESCC (n\u0026thinsp;=\u0026thinsp;50)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eMid-Ad ESCC (n\u0026thinsp;=\u0026thinsp;50)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003evalue\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eDemographic\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGender\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eMale\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e37\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eFemale\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eAge (years) (mean, range)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e58 (50\u0026ndash;70)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e60 (50\u0026ndash;70)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e59 (50\u0026ndash;70)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.15\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eBMI\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(mean, range)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e22.9\u003c/p\u003e\n\u003cp\u003e(15.2\u0026ndash;28.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e22.2\u003c/p\u003e\n\u003cp\u003e(17.6\u0026ndash;29.8)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e22.0\u003c/p\u003e\n\u003cp\u003e(16.4\u0026ndash;27.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.28\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSalted foods (like/dislike)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e13 (29)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e16 (34)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e25 (25)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.106\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSmoking (yes/no)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e22 (20)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e30 (20)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e31 (19)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.667\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eDrinking (yes/no)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20 (22)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20 (30)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e29 (21)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.201\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eFresh vegetables and fruits (like/dislike)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e30 (12)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e38 (12)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e34 (16)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.654\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eFamily history (with/without)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0 (42)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10 (40)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6 (44)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.005\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eHypertension (with/without)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7 (35)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e13 (37)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7 (43)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.282\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eDiabetes (with/without)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1 (41)*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6 (44)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5 (45)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.212\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eClinical\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"4\" align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eDifferentiation\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eG0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e27\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eG1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e16\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eG2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e21\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eG3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"8\" align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eTNM stages\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e27\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIB\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIIA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e19\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIIB\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIIIA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIIIB\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e19\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIVA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"5\" align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSurgery\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMATHE\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMcKeown\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e24\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSweet\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e25\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIvorLewis\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eESD\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003ctfoot\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"6\"\u003eNote: Patients were classified into 3 groups; HC, healthy controls; Mid-Ad, middle-advanced; ESCC, esophageal squamous cell carcinoma; MATHE: mediastinoscope-assisted transhiatal esophagectomy; ESD: endoscopic submucosal dissection; P value in chi-square test or t test. *There were 8 unknown findings.\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tfoot\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eAdditional 21 HC samples were acquired from a previously published study[\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e]. The HC cohort was identified as satisfying the study inclusion criteria, which were specifically negative for any form of cancer. Detailed pathological and clinical information was obtained from clinical records and illustrated in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and supplementary Table S1. The subject age and gender distributions among HC, Early ESCC and Mid-Ad ESCC groups were unbiased based on the Kruskal-Wallis H-test (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;0.11, supplementary Figure S1A) and Fisher test (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;\u0026gt;\u0026thinsp;0.5, supplementary Figure S1B), respectively. Our primary aim was to develop a convenient and excellent diagnostic model using new biotechnology reflecting genome-wide signatures from plasma cfDNA to distinguish patients with Early ESCC from HC subjects. For biomarker screening and classifier model construction for early diagnosis of ESCC, the ESCC patients and HC individuals were randomly divided into two groups: about 2/3 individuals as a training set and 1/3 individuals as an independent test set (also named as the internal test set) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and supplementary Figure S2). We also downloaded 5hmC-sequencing data from previously published studies of ESCC (as another independent external test set, also named as a Northern ESCC cohort since these ESCC patients were mainly from the North of China)[\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e\n\u003c/div\u003e\n\u003ch2\u003eConserved 5hmc Modification Changes In Escc, Potential Biomarkers For Diagnosis\u003c/h2\u003e\n\u003cp\u003eTo explore the distribution patterns of hydroxymethylation in plasma cfDNAs across the genome, we first identified 5hmC-enriched regions in each sample and subsequently determined as reliable peaks by high enrichment and significance as described in Methods. Comparison of global 5hmC modification levels between ESCC and HC groups revealed significant differences. The density of 5hmC peaks number were calculated and exhibited a broader distribution in ESCC group compared with the HC group that displayed a sharper and narrower curve (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eA). 5hmC peaks number in ESCC cfDNAs cohort was significantly more than that in HC cohort (Mann-Whitney U-test, \u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;3.11\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e, supplementary Figure S3) and showed a gradually increasing trend from stage 0 to stage IV patients (Mann-Kendall Test, \u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;1.65 \u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eB), indicating that 5hmC modification changes might play an important role in promoting tumorigenesis and progression of esophagus cancer. In addition, increased 5hmC modification levels within promoter and genebody regions in ESCC group were observed by 5hmC metagene profile analysis (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eC), consistent with previous observation in ESCC[\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e]. To further compare the cfDNA 5hmC modification difference between the two groups, we determined differential 5hmC peaks with DESeq2 package (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;\u0026lt;\u0026thinsp;0.01) and identified 398 5hmC up-regulation peaks and 227 5hmC down-regulation peaks in ESCC groups by comparing with HC groups. 5hmC up-regulation peaks from ESCC patients were significant enriched in promoters (28.14%) and 1st intron regions (15.58%) (i.e., mainly in the regulation regions of a gene) on the whole genome level, while more 5hmC down-regulation peaks were located in other introns (37.44%) and distal intergenic regions (39.21%) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eD).\u003c/p\u003e\n\u003cp\u003eNext, 5hmC motif enrichment analysis was performed in up-regulation and down-regulation regions of 5hmC signals in ESCC samples to further understand the correlation of 5hmC changes with potential interactions of binding proteins. The 5hmC up-regulation peaks in ESCC was significantly enriched in ERG motif (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1e-5, 28.83%), followed by ETS1 (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1e-4, 20.25%) and ETV2 motif (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1e-3, 17.79%), all of which belong to the large family of ETS transcription factors and bind to the consensus DNA sequence 5'-AGGAA-3' (left in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eE), most of which are downstream nuclear targets of Ras-MAP kinase signaling, and associate with cell development, differentiation, proliferation, apoptosis and tissue remodeling. The deregulation of ETS genes results in the malignant transformation of cells and is often observed in various types of malignant tumors[\u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e]. In contrast, the motif of GATA3 (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1e-5, 29.23%), GATA4 (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1e-5, 22.31%) and TRPS1 (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1e-4, 33.85%) were observed in 5hmC down-regulation peaks (right in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eE). GATA transcription factor mutations or dysregulated GATA expression were reported to associate with the wide range of diseases and pathologic phenotypes[\u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e]. These are consistent with previously published study showing that ETS motif and GATA motif were identified in 5hmC-gain and 5hmC-loss regions for esophageal cancer, respectively[\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e]. On the other hand, we further compared the enriched 5hmC motifs among our Southern ESCC cohort, Northern ESCC cohort[\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e] and colorectal cancer cohort[\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e] and observed. For the top 10-significantly hit 5hmC motifs, about 30.0%-50.0% 5hmC motifs shared between the Southern and the Northern ESCC cohorts, obviously more than those shared with CRC patients (about 10.0%-20.0%) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eF, supplementary Table S2). These results showed that the unique signature of plasma cfDNA 5hmC may represents a stable biomarker for discriminating ESCC and healthy individuals and could be potentially useful for early ESCC diagnosis.\u003c/p\u003e\n\u003ch2\u003eScreening, Validation And Performance Of Candidate 5hmc Biomarkers And Classifier\u003c/h2\u003e\n\u003cp\u003eSince the results of 5hmC up-regulation peaks were significantly enriched in promoter and 1st intron regions and 5hmC down-regulation peaks were significantly enriched in other introns and distal intergenic regions in the ESCC patients, which indicated that differential 5hmC signals in promoter and certain genebody regions might be utilized for identifying candidate 5hmC biomarkers to develop a convenient and stable diagnostic model. All 171 samples were randomly separated into two groups (54 HC, 34 Early ESCC and 31 Mid-Ad ESCC for training set, the remaining samples for internal test set) for development and evaluation of diagnostic model. The trained model was also validated on the independent previously-published Northern ESCC dataset (external test set) included 150 esophageal cancer and 183 HC samples[\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e], which was profiled for 5hmC using a nano-hmC-Seal method similar to the one used in our study.\u003c/p\u003e\n\u003cp\u003eFirstly, 925 candidate 5hmC marker genes derived from promoter and genebody regions were selected by Wilcoxon rank-sum test \u003cem\u003eP\u003c/em\u003e values\u0026thinsp;\u0026lt;\u0026thinsp;0.001 in the training set. Using Recursive feature elimination - Cross Validation (RFECV) approach, we further identified a disease-specific panel of 273 5hmC marker genes according to the contribution of each gene in the model (supplementary Table S3), and the distinct 5hmC landscapes in cfDNA showed apparent separation between ESCC and HC groups by unsupervised hierarchical clustering analysis (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eA). Similarly, the principal component analysis (PCA) also demonstrated distinct signatures that could discriminate the majority of the individuals of the two groups (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eB). To explore the biological significance of 273 marker genes with differential 5hmC signals, we performed gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analysis and found that 273 5hmC biomarkers were enriched in pathways associated with cancer and metastasis and mapped to tumor-related genes (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eC). For instance, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eD exhibited the IGV plot of the high-weight biomarker located at the FOXK1 gene, which plays an oncogenic role in the development of esophageal cancer[\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e]. The ESCC samples showed increased 5hmC modification level at FOXK1 gene when compared with the healthy controls. The classifier model based on 273 marker gene illustrated decent capacity for distinguishing ESCC patients from HC individuals in both internal test set Area under curve (AUC)\u0026thinsp;=\u0026thinsp;0.810 (95% CI: 0.693\u0026ndash;0.927); sensitivity\u0026thinsp;=\u0026thinsp;74.3%; specificity\u0026thinsp;=\u0026thinsp;82.4%) and external test set (AUC\u0026thinsp;=\u0026thinsp;0.862 (95% CI: 0.822\u0026ndash;0.902); sensitivity\u0026thinsp;=\u0026thinsp;69.3%; specificity\u0026thinsp;=\u0026thinsp;90.7%) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eE and \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eF). The prediction performance of classifier model in external test set was superior to that of internal test set, probably because of patients with 0 stage accounting for 27% (27/100) of our ESCC cohort who might be misclassified as HC individuals.\u003c/p\u003e\n\u003cp\u003eWe next assess the prediction accuracy of 5hmC biomarker classifier for different clinical stages in our ESCC patient subgroups. Box plots exhibited that the probability of being predicted as cancer gradually increasing with the progression of cancer stage in the test set (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eG). The 5hmC score between Early ESCC (stage 0 and I) and HC individuals demonstrated statistically significant disparity (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;4.35\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e, supplementary Figure S4), which suggested the capacity of the 5hmC model to discriminate Early ESCC patients from HC individuals. Meanwhile, 5hmC score also accurately distinguished between stage I and HC samples (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;4.54\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e), stage 0 and stage I, II, III-IV (\u003cem\u003eP\u003c/em\u003e values\u0026thinsp;=\u0026thinsp;4.64\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e, 1.46\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e and 9.85\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e, respectively, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eG). However, in differentiating stage 0 from HC samples, 5hmC score was not the best diagnostic feature (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;0.40, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eG), and the diagnostic accuracy (33.3%, supplementary Figure S5B) need to be further improved.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIntegrated model based on cfDNA signatures of low-pass WGS and 5hmC biomarkers improved diagnostic score for Early ESCC\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo explore the prediction potential of plasma cfDNA and search for more effective biomarkers, we employed low-pass whole-genome sequencing (WGS) to acquire genome-wide 5\u0026rsquo; end motif[\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e], nucleosome footprint (NF)[\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e] and fragmentation[\u003cspan class=\"CitationRef\"\u003e8\u003c/span\u003e] profiles from 71 HC and 93 ESCC samples with enough plasma cfDNA. Differential 5\u0026rsquo; end motif was identified by Wilcoxon rank-sum test \u003cem\u003eP\u003c/em\u003e value\u0026thinsp;\u0026lt;\u0026thinsp;0.001 and exhibited apparent separation for ESCC against HC samples by unsupervised hierarchical clustering analysis (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eA). Meanwhile, NF heatmap analysis indicated that genes with differential reads coverage between promoter and background regions (Wilcoxon rank-sum test, \u003cem\u003eP\u003c/em\u003e value\u0026thinsp;\u0026lt;\u0026thinsp;0.001) held power to distinguish ESCC from HC (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eB). Additionally, our data detected that the cfDNA fragment size of ESCC was more variable and much shorter (median size\u0026thinsp;\u0026lt;\u0026thinsp;150 base-pairs (bp)) than that of HC group (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eC). Collectively, all three genome features of cfDNA showed promising diagnostic potential for ESCC.\u003c/p\u003e\n\u003cp\u003eAs illustrated in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and supplementary Figure S2, healthy control individuals and patients with Early ESCC and Mid-Ad ESCC were randomly assigned into a training set (about 2/3 of samples, including 54 HC, 30 Early ESCC and 29 Mid-Ad ESCC) and a test set (about 1/3 of samples, including 17 HC, 15 Early ESCC and 19 Mid-Ad ESCC). Candidate biomarkers of the above three genomic features (5\u0026rsquo; end motif, NF and fragmentation) were screened to distinguish ESCC and HC in the training set. Least absolute shrinkage and selection operator (LASSO) regression method were applied to further reduce the number of candidate biomarkers. Eventually, 120 differential motif types, 170 differential NF genes and 10 fragment areas were selected for further analysis (supplementary Table S4A-S4C). Then, the Support Vector Machine (SVM) method was implemented for individual genomic feature-based model construction. The set of 120 motif types resulted in a discrimination model achieving an AUC value of 0.870 (95% CI: 0.769\u0026ndash;0.972) with sensitivity of 73.5% at specificity of 82.4% for ESCC patient classification in the test set (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eD, supplementary Figure S5A). The motif prediction model exhibited superior diagnostic power outperformed that of NF and fragmentation classifier model in cfDNA, which achieving an inferior performance with an AUC value of 0.813 (95% CI: 0.665\u0026ndash;0.961) and 0.806 (95% CI: 0.677\u0026ndash;0.936), respectively (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eD). Compared with the 5\u0026rsquo; end motif model, the fragmentation model demonstrated a higher sensitivity of 79.4% at the same specificity of 82.4%, and NF model showed excellent sensitivity of 91.2% but a lower specificity of 70.6% (supplementary Figure S5A). Because the above three genomic features and 5hmC depict different aspects of the genome, we envisioned that conjoint analysis would improve diagnostic power. Using the predictive score of the four individual models as input features, an integrated diagnostic model of combining low-pass WGS cfDNA signatures and 5hmC biomarkers was constructed. The diagnostic power of integrated model outperformed any individual genomic features, achieving an excellent AUC value of 0.934 (95% CI: 0.867-1.000) with a sensitivity of 82.4%, specificity of 88.2% and accuracy of 84.3% for ESCC patient classification in the test set (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eD, \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eE).\u003c/p\u003e\n\u003cp\u003eThe combine scores showed an increasing trend from HC to ESCC, noting the significantly higher scores in ESCC patients with stage 0 and stage I than in subjects of HC in the test set (\u003cem\u003eP\u003c/em\u003e values\u0026thinsp;=\u0026thinsp;1.10\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e and 3.31\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e, respectively, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eF), reinforcing the idea that integrated model had great potentials as a new strategy for ESCC early diagnosis and surveillance. The integrated classifier model had good but slightly reduced power to call stage III-IV patients (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eF), who in general suffered from metastasis to various tissues and are expected to have more complex tumorous DNA profiles. In addition, we compared the diagnostic performance between 5hmC model and integrated model on different clinical stage of ESCC samples in test set (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eG). For Early ESCC patients, especially patients with stage 0 who have low-disease burden, the performance of integrated model was obviously superior to that of 5hmC model (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eG), showing evidently higher prediction accuracy (80.0% vs 33.3%) (supplementary Figure S5B). Integrated model and 5hmC model validated similar performance in stage II patients. Instead, for advanced ESCC patients (stage IIIB and IVA), 5hmC model showed an improved diagnostic performance compared to integrated model. These data demonstrated that genome-wide integration was a sensitive and robust approach which could perform better than 5hmC-based methods for accurate diagnosis of early stage esophageal carcinoma.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eWe ultimately constructed an integrated classifier using low-pass WGS and 5hmC biomarkers to realize our primal goal for early detection of esophageal cancer which may bring convenience to clinical practice. At first step, an investigation of the enriched 5hmC motifs in genome was conducted, and the outcomes showed that these existing some enriched 5hmC motifs in different regions like ERG, ETS1 and ETV2 which were proved to associate with cell developments, differentiation, proliferation, apoptosis and tissue remodeling, as well as GATA3, GATA4 and TRPS1 had been reported to be connected with a wide range of diseases and pathologic phenotypes, and both of these findings were consistent with our previous awareness about the biological behavior of tumors. We then compared the enriched 5hmC motifs among our Southern ESCC cohort (with 150 pb paired-end sequencing mode), the Northern ESCC cohort (with 50 bp paired-end sequencing mode)[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] and colorectal cancer cohort (with 150 pb paired-end sequencing mode)[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] which suggested that some 5hmC motif signal changes were more conservative in ESCC patients than in other patients (such as CRC), which were not effected by different sequencing platforms. These ESCC-unique 5hmC motifs mainly consisted of up-regulated ERG, ETS1 and ETV2 of ETS family, which were reported as cancer-associated transcription factors regulating Ras-MAP kinase signaling pathways[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan additionalcitationids=\"CR33 CR34 CR35 CR36\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. Up-regulated EHF and ELF3 of ETS family uniquely consisted in the Northern ESCC 5hmC motifs, indicating the bias possibly due to geographic distribution or different pathological subtype compositions between these two ESCC cohorts. Other functions, such as vasculature development, muscle structure development, cell part morphogenesis/development and hippo signaling pathway were also consistent with those reported in the Northern ESCC cohort, indicating their essential roles in ESCC tumorigenesis and developments. In addition, other different functions among 5hmC biomarkers between our study and the Northern ESCC cohort might reveal different molecular mechanisms in different pathological subgroups of ESCC cohort or different epigenetic modifications resulting from different living habitats, which needed further investigation on large scale population studies.\u003c/p\u003e \u003cp\u003eRecently, some relevant studies showed somewhat dissatisfactory early detection efficiency which results were around 50% and 62.5% in sensitivity on esophageal cancer stage 0 and I respectively using cfDNA methylation method[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e], and some of the others exhibited a good sensitivity of\u003c/p\u003e \u003cp\u003e93.75% and specificity of 85.71% (AUC\u0026thinsp;=\u0026thinsp;0.972) in distinguishing ESCC from HC groups but slightly ignored the discussion focusing on the early stage using 5hmC method[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. In the present study, using our selected 273 5hmC marker genes, we indeed succeeded in distinguishing ESCC patients from HC individuals in both internal test set. Besides, we also made a further analysis to explore the prediction abilities of 5hmC biomarkers classifier in different stages in ESCC subgroups and found that there were significant differences between Early ESCC (stage 0 and I) and HC individuals, and the differences between stage I and HC samples (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;4.54\u0026times;10\u0026thinsp;\u0026minus;\u0026thinsp;3), stage 0 and stage I, II, III-IV were statistically significant too All the outcomes showed that the 5hmC indeed played an important role in tumor progress. Regretfully, in distinguishing stage 0 from HC, we failed to work out a precise identification (\u003cem\u003eP\u003c/em\u003e value\u0026thinsp;=\u0026thinsp;0.40, with accuracy of 33.3%) which needed further investigation.\u003c/p\u003e \u003cp\u003eSo far, we had verified that the 5hmC could definitely serve as a favorable and non-invasive tool in differential diagnosis of esophageal cancer However, with the consideration of the importance of early detection of esophageal cancer and the deficient detection efficiency between early-stage esophageal cancer (stage 0) and HC samples in our study, we aim to establish a novel diagnostic model which can achieve our goals for early detection to satisfy clinical real needs. After employing low-pass whole-genome sequencing(WGS) technology, we worked out an integrated diagnostic classifier consisting of genome-wide 5\u0026rsquo; end motif, nucleosome footprint (NF), fragmentation profiles and 5hmC biomarkers, and eventually realized an excellent AUC value of 0.934 with a sensitivity of 82.4%, specificity of 88.2% and accuracy of 84.3% for ESCC patient classification in the test set, especially the scores in ESCC patients with stage 0 and 1 were significantly higher than that in HC subjects(\u003cem\u003eP\u003c/em\u003e values\u0026thinsp;=\u0026thinsp;1.10\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e and 3.31\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e, respectively). Briefly speaking, our findings of combination of low-pass WGS of cfDNA signatures and 5hmC biomarkers improved the classifier\u0026rsquo;s efficiency from 65\u0026ndash;82% of sensitivity at the specificity of 88% on an overall level. Most importantly, for stage 0 patients who had low-disease burden, the combined classifier significantly improved the prediction accuracy from 33.3\u0026ndash;80.0%. Though the test cohort population in this study was still limited, it was still a valuable attempt for using combined approaches to improve prediction accuracy in non-invasion early diagnosis of ESCC.\u003c/p\u003e \u003cp\u003eIn general, more and more studies showed that integrating multi-omics detection was a promising methodology for non-invasive early diagnosis of many types of cancer. 5hmC was an intermediate produced through oxidation of 5-methylcytosine (5mC) by the oxygenase ten-eleven translocation (TET) enzyme family during 5mC demethylation process[\u003cspan additionalcitationids=\"CR40\" citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Both 5mC and 5hmC were presumed to have an important role in gene expression and regulation, and their modification changes were both observed in a wide range of malignant tumors, including ESCC[\u003cspan additionalcitationids=\"CR43 CR44 CR45\" citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. Therefore, from a diagnostic perspective, combination of 5mC and 5hmC signal detection, as well as WGS features of cfDNA in this study to construct classifiers to further improve the specificity and sensitivity for early diagnosis of different subtypes and stages of ESCC patients was worthy of attempts. Here, an integrated model was worked out to realize early detection for esophageal cancer in a non-invasive approach that could be conveniently applied to clinical practices. Of course, further investigations about the stability of this model, the discriminating capabilities for different subtypes of esophageal cancer, or the practical application values are needed to execute. We sincerely hope that this study could provide an innovative clinical diagnostic strategy and could be put into practical use in the future which ultimately bring our patients with positive benefits.\u003c/p\u003e"},{"header":"Conclusions","content":" \u003cp\u003eThis study suggests that the blood-based 5hmC integrated with low-pass WGS model could improve the accurate diagnosis of early stage ESCC, particularly for diagnosis of very early stage ESCC.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eStudy participants and clinical features\u003c/h2\u003e \u003cp\u003eAll the patients selected to experimental group were diagnosed with esophageal squamous cell carcinoma (ESCC) and were confirmed cytopathologically and histologically, also, the patients were restricted to which on initial treatments from 0-PIV diagnosed by using the esophagus and esophagogastric junction of the eighth edition of the AJCC/UICC cancer staging manuals[\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. Except for living habits (especially hot food preference) and family disease history, other hazard covariates like BMI, smoking and drinking were kept consistent to the greatest extent. For patients subjected to control group, participants were selected from health examine center of our hospital, and no esophageal squamous cell carcinoma and other relevant diseases were observed. The other standards were consistent with experimental group. We excluded patients who used to receive surgery, chemoradiotherapy or immunotherapy, and patients who suffered from other illnesses like leukemia, neurodegenerative disease or other tumoral disorders. In total, 150 participants consisting of 100 esophageal squamous cell carcinoma (ESCC) patients and 50 healthy controls (HC) were enrolled from Nanfang Hospital of Southern Medical University in March 2019 to December 2020 and named as a Southern ESCC cohort since the patients were mainly from the Southern of China. Moreover, we added additional 21 healthy control samples used in a previously published study[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] to our HC cohort.\u003c/p\u003e \u003c/div\u003e\n\u003ch2\u003eBlood Sample Preparation And Cfdna Extraction\u003c/h2\u003e\n\u003cp\u003ePeripheral blood specimens (10 mL/subject) were obtained from patients who were newly diagnosed and had never received any medication or radical treatment for disease, which was prior to any biopsy or surgical resection. 71 HC blood samples were collected at the time of visiting the clinic for routine physical examination. All peripheral blood samples were stored in cell-free tubes (Streck, USA) at 4\u0026deg;C for no more than 72 hours before being separated into plasma and stored at -80\u0026deg;C in the laboratory. The plasma cell free DNA (cfDNA) was isolated using the MagMAX Cell-Free DNA Isolation Kit (Thermo, USA) according to the manufacturer's protocol. The quality of purified DNA was quantified by Qubit\u0026reg; 4.0 Fluorometer (Life Technologies, USA), and the DNA fragment size composition was assayed by Fragment Analyzer (Agilent, USA).\u003c/p\u003e\n\u003ch2\u003e5hmc Sequencing And Data Processing\u003c/h2\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e5hmC library construction and sequencing\u003c/h2\u003e \u003cp\u003e5hmC library construction was performed according to the method previously described[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Briefly, the purified cfDNA (5\u0026ndash;20 ng) were end-repaired and A tailed (5X ER/A-Tailing Enzyme Mix, Enzymatics, USA) then ligated with T-adaptors on both ends (WGS Ligase, Enzymatics, USA), to result in pre-library. Subsequently, ligated DNA was incubated in a 25 \u0026micro;l solution containing 50 mM HEPES buffer (pH\u0026thinsp;=\u0026thinsp;8.0), 25 mM MgCl2, 60 \u0026micro;M UDP-6-N3-Glc (Active Motif, USA) and 12.5 U βGT (Thermo, USA) for 2 h at 37\u0026deg;C. Then, 2.5 \u0026micro;l DBCO-PEG4-biotin (Click Chemistry Tools, USA) was added to the reaction mixture and incubated for 2 h at 37\u0026deg;C. The purified DNA was incubated with 0.5 \u0026micro;l M270 Streptavidin beads (Life Technologies, USA) pre-blocked with salmon sperm DNA in buffer 1 (5 mM Tris pH 7.5, 0.5 mM EDTA, 1 M NaCl and 0.2% Tween 20) for 30 min. Afterwards, DNA fragments containing 5hmC features were subjected to PCR amplification, followed by the purification of the PCR products using AMPure XP beads according to the manufacturer's instructions. Finally, sequenced using Illumina NovaSeq 6000 platform.\u003c/p\u003e \u003c/div\u003e\n\u003ch2\u003eMapping And Sequencing Quality Control \u003c/h2\u003e\n\u003cp\u003eThe raw sequencing reads were removed adaptor and end sequence by trim_galore software (https: //\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003ca href=\"http://www.bioinformatics\" target=\"_blank\"\u003ewww.bioinformatics\u003c/a\u003e\u003c/span\u003e\u003c/span\u003e. babraham. ac.uk/projects/trim_galore/)[\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. Acquired clean data were aligned to the human reference genome (hg19/GRCh37) by Bowtie2 v2.2.5 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://bowtiebio.sourceforge.net/bowtie2/index.shtml\u003c/span\u003e\u003c/span\u003e)[\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]. Picard Tools (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://broadinstitute.github.io/picard/\u003c/span\u003e\u003c/span\u003e) and SAMtools (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://samtools.sourceforge.net/\u003c/span\u003e\u003c/span\u003e)[\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e] were used to process and filter PCR duplicates for mapped BAM files. Reads with duplicate ratio less than 65% and enrichment efficiency over 95-fold passed the quality control and were used for further analysis.\u003c/p\u003e\n\u003ch2\u003e5hmc Peak Identification\u003c/h2\u003e\n\u003cp\u003eModel-based analysis of ChIP-seq[\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e] was used to identify the 5hmC-enriched regions in each sample (the q value cut-off to call significant regions is 0.01; model fold = [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]). The peaks with high enrichment and significance (q\u0026thinsp;\u0026lt;\u0026thinsp;1E-12; fold enrichment\u0026thinsp;\u0026gt;\u0026thinsp;8) in all samples were considered as highly reliable 5hmC-enriched peaks. The 5hmC enrichment level was expressed as fragments per kilobase of 5hmC-DNA per million fragments mapped (FPKM). The genomic annotation of 5hmC peak regions was performed using annotatr[\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e] and the genome-wide distribution of 5hmC was visualized using the Integrated Genomics Viewer[\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e, \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. The metagene profile was generated using ngsplot[\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e].\u003c/p\u003e\n\u003ch2\u003eDifferential 5hmc Peak Regions Detection\u003c/h2\u003e\n\u003cp\u003eThe differential 5hmC peak regions between healthy control group and ESCC group were identified using DESeq2 package[\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e] with \u003cem\u003eP\u003c/em\u003e value\u0026thinsp;\u0026lt;\u0026thinsp;0.01. De novo motif analysis around differential 5hmC peaks was performed using HOMER software (version 4.9). Functional gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses were performed using an online tool of metascape (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://metascape.org/\u003c/span\u003e\u003c/span\u003e)[\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e].\u003c/p\u003e\n\u003ch2\u003e5hmc Biomarkers Identification And Evaluating Performance\u003c/h2\u003e\n\u003cp\u003e5hmC candidate biomarkers for cancer prediction models were identified firstly based on Wilcoxon rank-sum test (\u003cem\u003eP\u003c/em\u003e values\u0026thinsp;\u0026lt;\u0026thinsp;0.001) between ESCC and HC groups and then processed to reduce the number of 5hmC biomarkers based on Recursive Feature Elimination - Cross Validation (RFECV) approach in the training set (54 HC, 34 Early ESCC and 31 Mid-Ad ESCC). The remaining samples in each group were used as the internal test set (including 17 HC, 16 Early ESCC and 19 Mid-Ad ESCC). In addition, we downloaded 150 esophageal cancer plasma-5hmC data and 183 healthy control plasma-5hmC data from the article previously published (nominated as the Northern ESCC cohort)[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] as an external test set to validate our results.\u003c/p\u003e \u003cp\u003eThe selected 273 differential 5hmC biomarkers using for sample identification to be ESCC or HC were analyzed by the principal component analysis (PCA). A heatmap of R package (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cran.r-project.org/web/packages/pheatmap/\u003c/span\u003e\u003c/span\u003e index.html) was used to visualize hierarchical clustering and the distance in a heatmap figure[\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. 5hmC biomarkers were further processed for classification model construction based on a supervised two-class Support Vector Machine (SVM) method. To optimize the classification model and ensure the significance of each potential 5hmC biomarkers, the GridSearchCV package in Python in conjunction with cross-validation were performed to obtain the optimal parameters for SVM based on the Gaussian kernel (kernel='rbf', gamma\u0026thinsp;=\u0026thinsp;0.001, the penalty parameter C\u0026thinsp;=\u0026thinsp;4). The defined 5hmC-DNA regions and their corresponding genes were finally applied to classify the test set samples.\u003c/p\u003e\n\u003ch2\u003eLow-pass Whole Genome Sequencing And Data Processing\u003c/h2\u003e\n\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eWGS library construction and sequencing\u003c/h2\u003e \u003cp\u003e1\u0026ndash;10 ng cfDNA were end-repaired and A-tailed (Berry, China) then ligated with T-adaptors (Berry, China), to result in pre-library. The pre-libraries were purified by Clean NGS beads (VdoBiotech, China), followed by being quantified by the KAPA Library Quantification Kit (Kapa Biosystems, USA). cfDNA fragment size was confirmed using Bioanalyzer (Agilent, USA). Sequencing libraries were pooled at equal amount and then sequenced on an Illumina CN500 platform (Illumina, San Diego, USA) with an average coverage of 2x.\u003c/p\u003e \u003c/div\u003e\n\u003ch2\u003eMapping And Sequencing Quality Control\u003c/h2\u003e\n\u003cp\u003eThe raw sequencing reads were removed adaptor and end sequence by fastp software (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/OpenGene/fastp\u003c/span\u003e\u003c/span\u003e). Acquired clean data were aligned to human reference genome (hg19/GRCh37) using bwa-mem (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/lh3/bwa\u003c/span\u003e\u003c/span\u003e). Duplicate reads were marked by sambamba (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/biod/sambamba/\u003c/span\u003e\u003c/span\u003e). Sequencing data were further processed by SAMtools (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://samtools.sourceforge.net/\u003c/span\u003e\u003c/span\u003e)[\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e] to get rid of marked duplicates, unmapped reads and low quality reads. Reads with duplicate rate less than 15% and mapping rate more than 95% passed the quality control and were used for further analysis.\u003c/p\u003e\n\u003ch2\u003eWgs-based Biomarkers Identification And Integrated Model Construction\u003c/h2\u003e\n\u003cp\u003eTo select more effective biomarkers for distinguishing ESCC samples from healthy controls, all samples were randomly separated into two subsets: the training set consisted of 54 HC, 30 Early ESCC and 29 Mid-Ad ESCC, and the test set consisted of the remaining samples (including 17 HC, 15 Early ESCC and 19 Mid-Ad ESCC). The Wilcoxon rank-sum test was used to compare biomarker features between ESCC and HC groups. Least Absolute Shrinkage and Selection Operator (LASSO) methods were applied to further reduce the number of biomarkers in the training set. The detailed selecting process was performed as follows.\u003c/p\u003e \u003cp\u003e \u003cb\u003e5\u0026rsquo; end Motif\u003c/b\u003e: 256 different types of 4mer 5\u0026rsquo; end motif were identified and calculated their percentages (using pysam (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pysam.readthedocs.io/en/latest/\u003c/span\u003e\u003c/span\u003e)) without considering chromosome Y and unidentifiable bases. Following motif types were filtered out: 1) \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026ge;\u0026thinsp;0.05 in Wilcoxon rank-sum test between ESCC and HC groups; 2) weight of 0 via LASSO. Eventually, 120 motif types were left for further analysis.\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eNucleosome footprint (NF)\u003c/strong\u003e \u003cp\u003eWe obtained all transcripts of coding genes, microRNAs and long non-coding RNAs (LncRNAs) and calculated the distance between transcripts. The transcripts with the distance more than 200 bp were retained. If the distance between two transcripts less than 200 bp, the longer transcript was retained. A total of 57151 transcripts of 30588 genes were recruited for analysis. The promoter region and background region of transcripts were divided, and the reads number of different regions was counted with featureCounts[\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]. NF score of each gene is calculated as\u003c/p\u003e \u003c/p\u003e \u003cp\u003eNF Score=(background1\u0026thinsp;+\u0026thinsp;background2)/2-Promotor\u003c/p\u003e \u003cp\u003eFollowing genes were filtered out: 1) more than 10% of the total samples showing an NF score of 0; 2) \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026ge;\u0026thinsp;0.001 in Wilcoxon rank-sum test between ESCC and HC groups; 3) weight of 0 via LASSO. Eventually, 170 genes were left for further analysis.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFragment\u003c/b\u003e: The whole genome except Y chromosome was divided into 1M sized bins, resulting in 3055 areas. Pysam (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pysam.readthedocs.io/en/latest/\u003c/span\u003e\u003c/span\u003e) was used to calculate the length of insertion fragment and ratio of short/long fragment in different regions. LASSO was then used to filter out areas with a weight of 0, and finally 10 areas were retained.\u003c/p\u003e \u003cp\u003eThereafter, the Support Vector Machine (SVM) method was implemented for individual genomic feature-based model construction. 10-fold cross-validation method was employed to optimize the combination of the parameters in the training set, and cut-off value was set at the point with the best diagnostic accuracy. To obtain the best diagnostic model, logistic regression model was generated using the predictive score of the four individual models as input features to integrate the outcome of each model based on the training dataset. The logistic Score was calculated as follows.\u003c/p\u003e \u003cp\u003eLogistic Score\u0026thinsp;=\u0026thinsp;exp(Z) / (1\u0026thinsp;+\u0026thinsp;exp(Z)), where Z= -2.57+(3.35\u0026times;NF)+༈0.05\u0026times;Fragment༉+༈0.75\u0026times;Motif༉+༈1.74\u0026times;5hmC༉\u003c/p\u003e \u003cp\u003eReceiver operating characteristic (ROC) curves[\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e] were generated to evaluate the performance of a prediction algorithm, using the pROC[\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e] library in the R package. Sensitivity and specificity were estimated at the score cut-off that maximizes the sum of sensitivity and specificity using the ROCR library in the R package.\u003c/p\u003e\n\u003ch2\u003eStatistics\u003c/h2\u003e\n\u003cp\u003eThe statistical methods used were stated in the above methods. The R code related to classifier detection and modelling is available upon request.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study was approved by the ethics committee of the Nanfang Hospital, Southern Medical University, Guangzhou, China (reference: NFEC-2019-014) and was also registered with ClinicalTrials.gov (reference: NCT03922230). Besides, this study was conducted by the approval of the Institutional Review Board of Nanfang Hospital of Southern Medical University, and the written informed consents were obtained from all participants according to the institutional guidelines.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll the authors agreed to submit and publish the manuscript to Journal of Hematology \u0026amp; Oncology.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll of the raw and processed data used in this study have been uploaded to the Genome Sequence Archive depository (https://ngdc.cncb.ac.cn/gsa)[\u003ca href=\"#_ENREF_62\"\u003e62\u003c/a\u003e] with the Accession Number (HRA001476).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest:\u003c/strong\u003eSW, HL, FS, SW and XZ are employees of Berry Oncology Corporation.Other authors had no declaration of conflicts of interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e. Science and Technology Planning Project of Guangdong Province(2017B020226005).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eKC, SW and DLdesigned the research. DL,XW, JZ and XD recruited the examinees. DL, XW and SF got consents and collected raw data. DL, XY, HL and SWwrote the manuscript. SW, XZ,FS conducted bioinformatics analysisanddata visualization. All authors reviewed and approved the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThanks for the help of Professor Side Liu, Dr. Jianqun Cai and Dr. Jing Wang from Department of Gastroenterology, Ms. Li Zhen from Department of General Surgery, Ms. Yu Guo from Department of Huiqiao Building, and Ms. Mei Li from Department of Thoracic Surgery, Nanfang Hospital, Southern Medical University, Guangzhou, China. This study was funded by the Science and Technology Planning Project of Guangdong Province (2017B020226005).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F: \u003cstrong\u003eGlobal Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries\u003c/strong\u003e. \u003cem\u003eCA: a cancer journal for clinicians \u003c/em\u003e2021, \u003cstrong\u003e71\u003c/strong\u003e(3):209-249.\u003c/li\u003e\n\u003cli\u003eSmyth EC, Lagergren J, Fitzgerald RC, Lordick F, Shah MA, Lagergren P, Cunningham D: \u003cstrong\u003eOesophageal cancer\u003c/strong\u003e. \u003cem\u003eNature reviews Disease primers \u003c/em\u003e2017, \u003cstrong\u003e3\u003c/strong\u003e:17048.\u003c/li\u003e\n\u003cli\u003eAllemani C, Matsuda T, Di Carlo V, Harewood R, Matz M, Nik\u0026scaron;ić M, Bonaventure A, Valkov M, Johnson CJ, Est\u0026egrave;ve J\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGlobal surveillance of trends in cancer survival 2000-14 (CONCORD-3): analysis of individual records for 37\u003c/strong\u003e\u003cstrong\u003e 513\u003c/strong\u003e\u003cstrong\u003e 025 patients diagnosed with one of 18 cancers from 322 population-based registries in 71 countries\u003c/strong\u003e. \u003cem\u003eLancet \u003c/em\u003e2018, \u003cstrong\u003e391\u003c/strong\u003e(10125):1023-1075.\u003c/li\u003e\n\u003cli\u003eWani S, Yadlapati R, Singh S, Sawas T, Katzka DA, Hall M, Bergman J, Canto MI, Chak A, Corley DA\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003ePost-Endoscopy Esophageal Neoplasia in Barrett\u0026rsquo;s Esophagus: Consensus Statements from an International Expert Panel\u003c/strong\u003e. \u003cem\u003eGastroenterology \u003c/em\u003e2021.\u003c/li\u003e\n\u003cli\u003edi Pietro M, Canto MI, Fitzgerald RC: \u003cstrong\u003eEndoscopic Management of Early Adenocarcinoma and Squamous Cell Carcinoma of the Esophagus: Screening, Diagnosis, and Therapy\u003c/strong\u003e. \u003cem\u003eGastroenterology \u003c/em\u003e2018, \u003cstrong\u003e154\u003c/strong\u003e(2):421-436.\u003c/li\u003e\n\u003cli\u003eWan JCM, Massie C, Garcia-Corbacho J, Mouliere F, Brenton JD, Caldas C, Pacey S, Baird R, Rosenfeld N: \u003cstrong\u003eLiquid biopsies come of age: towards implementation of circulating tumour DNA\u003c/strong\u003e. \u003cem\u003eNature reviews Cancer \u003c/em\u003e2017, \u003cstrong\u003e17\u003c/strong\u003e(4):223-238.\u003c/li\u003e\n\u003cli\u003eThierry AR, El Messaoudi S, Gahan PB, Anker P, Stroun M: \u003cstrong\u003eOrigins, structures, and functions of circulating DNA in oncology\u003c/strong\u003e. \u003cem\u003eCancer metastasis reviews \u003c/em\u003e2016, \u003cstrong\u003e35\u003c/strong\u003e(3):347-376.\u003c/li\u003e\n\u003cli\u003eCristiano S, Leal A, Phallen J, Fiksel J, Adleff V, Bruhm DC, Jensen S\u0026Oslash;, Medina JE, Hruban C, White JR\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGenome-wide cell-free DNA fragmentation in patients with cancer\u003c/strong\u003e. \u003cem\u003eNature \u003c/em\u003e2019, \u003cstrong\u003e570\u003c/strong\u003e(7761):385-389.\u003c/li\u003e\n\u003cli\u003eLuo H, Wei W, Ye Z, Zheng J, Xu RH: \u003cstrong\u003eLiquid Biopsy of Methylation Biomarkers in Cell-Free DNA\u003c/strong\u003e. \u003cem\u003eTrends in molecular medicine \u003c/em\u003e2021, \u003cstrong\u003e27\u003c/strong\u003e(5):482-500.\u003c/li\u003e\n\u003cli\u003eJones PA: \u003cstrong\u003eFunctions of DNA methylation: islands, start sites, gene bodies and beyond\u003c/strong\u003e. \u003cem\u003eNature reviews Genetics \u003c/em\u003e2012, \u003cstrong\u003e13\u003c/strong\u003e(7):484-492.\u003c/li\u003e\n\u003cli\u003eBaylin SB, Jones PA: \u003cstrong\u003eA decade of exploring the cancer epigenome - biological and translational implications\u003c/strong\u003e. \u003cem\u003eNature reviews Cancer \u003c/em\u003e2011, \u003cstrong\u003e11\u003c/strong\u003e(10):726-734.\u003c/li\u003e\n\u003cli\u003eDor Y, Cedar H: \u003cstrong\u003ePrinciples of DNA methylation and their implications for biology and medicine\u003c/strong\u003e. \u003cem\u003eThe Lancet \u003c/em\u003e2018, \u003cstrong\u003e392\u003c/strong\u003e(10149):777-786.\u003c/li\u003e\n\u003cli\u003eBranco MR, Ficz G, Reik W: \u003cstrong\u003eUncovering the role of 5-hydroxymethylcytosine in the epigenome\u003c/strong\u003e. \u003cem\u003eNature reviews Genetics \u003c/em\u003e2011, \u003cstrong\u003e13\u003c/strong\u003e(1):7-13.\u003c/li\u003e\n\u003cli\u003eGuler GD, Ning Y, Ku CJ, Phillips T, McCarthy E, Ellison CK, Bergamaschi A, Collin F, Lloyd P, Scott A\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eDetection of early stage pancreatic cancer using 5-hydroxymethylcytosine signatures in circulating cell free DNA\u003c/strong\u003e. \u003cem\u003eNature communications \u003c/em\u003e2020, \u003cstrong\u003e11\u003c/strong\u003e(1):5270.\u003c/li\u003e\n\u003cli\u003eZhang J, Han X, Gao C, Xing Y, Qi Z, Liu R, Wang Y, Zhang X, Yang YG, Li X\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e5-Hydroxymethylome in Circulating Cell-free DNA as A Potential Biomarker for Non-small-cell Lung Cancer\u003c/strong\u003e. \u003cem\u003eGenomics, proteomics \u0026amp; bioinformatics \u003c/em\u003e2018, \u003cstrong\u003e16\u003c/strong\u003e(3):187-199.\u003c/li\u003e\n\u003cli\u003eGilat N, Tabachnik T, Shwartz A, Shahal T, Torchinsky D, Michaeli Y, Nifker G, Zirkin S, Ebenstein Y: \u003cstrong\u003eSingle-molecule quantification of 5-hydroxymethylcytosine for diagnosis of blood and colon cancers\u003c/strong\u003e. \u003cem\u003eClinical epigenetics \u003c/em\u003e2017, \u003cstrong\u003e9\u003c/strong\u003e:70.\u003c/li\u003e\n\u003cli\u003eLi W, Zhang X, Lu X, You L, Song Y, Luo Z, Zhang J, Nie J, Zheng W, Xu D\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e5-Hydroxymethylcytosine signatures in circulating cell-free DNA as diagnostic biomarkers for human cancers\u003c/strong\u003e. \u003cem\u003eCell research \u003c/em\u003e2017, \u003cstrong\u003e27\u003c/strong\u003e(10):1243-1257.\u003c/li\u003e\n\u003cli\u003eSong CX, Yin S, Ma L, Wheeler A, Chen Y, Zhang Y, Liu B, Xiong J, Zhang W, Hu J\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages\u003c/strong\u003e. \u003cem\u003eCell research \u003c/em\u003e2017, \u003cstrong\u003e27\u003c/strong\u003e(10):1231-1242.\u003c/li\u003e\n\u003cli\u003eCai J, Chen L, Zhang Z, Zhang X, Lu X, Liu W, Shi G, Ge Y, Gao P, Yang Y\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGenome-wide mapping of 5-hydroxymethylcytosines in circulating cell-free DNA as a non-invasive approach for early detection of hepatocellular carcinoma\u003c/strong\u003e. \u003cem\u003eGut \u003c/em\u003e2019, \u003cstrong\u003e68\u003c/strong\u003e(12):2195-2205.\u003c/li\u003e\n\u003cli\u003eZviran A, Schulman RC, Shah M, Hill STK, Deochand S, Khamnei CC, Maloney D, Patel K, Liao W, Widman AJ\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGenome-wide cell-free DNA mutational integration enables ultra-sensitive cancer monitoring\u003c/strong\u003e. \u003cem\u003eNat Med \u003c/em\u003e2020, \u003cstrong\u003e26\u003c/strong\u003e(7):1114-1124.\u003c/li\u003e\n\u003cli\u003eChabon JJ, Hamilton EG, Kurtz DM, Esfahani MS, Moding EJ, Stehr H, Schroers-Martin J, Nabet BY, Chen B, Chaudhuri AA\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eIntegrating genomic features for non-invasive early lung cancer detection\u003c/strong\u003e. \u003cem\u003eNature \u003c/em\u003e2020, \u003cstrong\u003e580\u003c/strong\u003e(7802):245-251.\u003c/li\u003e\n\u003cli\u003eMathios D, Johansen JS, Cristiano S, Medina JE, Phallen J, Larsen KR, Bruhm DC, Niknafs N, Ferreira L, Adleff V\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eDetection and characterization of lung cancer using cell-free DNA fragmentomes\u003c/strong\u003e. \u003cem\u003eNature communications \u003c/em\u003e2021, \u003cstrong\u003e12\u003c/strong\u003e(1):5060.\u003c/li\u003e\n\u003cli\u003eChen L, Abou-Alfa GK, Zheng B, Liu JF, Bai J, Du LT, Qian YS, Fan R, Liu XL, Wu L\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGenome-scale profiling of circulating cell-free DNA signatures for early detection of hepatocellular carcinoma in cirrhotic patients\u003c/strong\u003e. \u003cem\u003eCell research \u003c/em\u003e2021, \u003cstrong\u003e31\u003c/strong\u003e(5):589-592.\u003c/li\u003e\n\u003cli\u003eSzymanski JJ, Sundby RT, Jones PA, Srihari D, Earland N, Harris PK, Feng W, Qaium F, Lei H, Roberts D\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eCell-free DNA ultra-low-pass whole genome sequencing to distinguish malignant peripheral nerve sheath tumor (MPNST) from its benign precursor lesion: A cross-sectional study\u003c/strong\u003e. \u003cem\u003ePLoS Med \u003c/em\u003e2021, \u003cstrong\u003e18\u003c/strong\u003e(8):e1003734.\u003c/li\u003e\n\u003cli\u003eXiao Z, Wu W, Wu C, Li M, Sun F, Zheng L, Liu G, Li X, Yun Z, Tang J\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e5-Hydroxymethylcytosine signature in circulating cell-free DNA as a potential diagnostic factor for early-stage colorectal cancer and precancerous adenoma\u003c/strong\u003e. \u003cem\u003eMolecular oncology \u003c/em\u003e2021, \u003cstrong\u003e15\u003c/strong\u003e(1):138-150.\u003c/li\u003e\n\u003cli\u003eTian X, Sun B, Chen C, Gao C, Zhang J, Lu X, Wang L, Li X, Xing Y, Liu R\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eCirculating tumor DNA 5-hydroxymethylcytosine as a novel diagnostic biomarker for esophageal cancer\u003c/strong\u003e. \u003cem\u003eCell research \u003c/em\u003e2018, \u003cstrong\u003e28\u003c/strong\u003e(5):597-600.\u003c/li\u003e\n\u003cli\u003eOikawa T: \u003cstrong\u003eETS transcription factors: possible targets for cancer therapy\u003c/strong\u003e. \u003cem\u003eCancer science \u003c/em\u003e2004, \u003cstrong\u003e95\u003c/strong\u003e(8):626-633.\u003c/li\u003e\n\u003cli\u003eRomano O, Miccio A: \u003cstrong\u003eGATA factor transcriptional activity: Insights from genome-wide binding profiles\u003c/strong\u003e. \u003cem\u003eIUBMB Life \u003c/em\u003e2020, \u003cstrong\u003e72\u003c/strong\u003e(1):10-26.\u003c/li\u003e\n\u003cli\u003eChen D, Wang K, Li X, Jiang M, Ni L, Xu B, Chu Y, Wang W, Wang H, Kang H\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eFOXK1 plays an oncogenic role in the development of esophageal cancer\u003c/strong\u003e. \u003cem\u003eBiochem Biophys Res Commun \u003c/em\u003e2017, \u003cstrong\u003e494\u003c/strong\u003e(1-2):88-94.\u003c/li\u003e\n\u003cli\u003eSerpas L, Chan RWY, Jiang P, Ni M, Sun K, Rashidfarrokhi A, Soni C, Sisirak V, Lee W-S, Cheng SH\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003edeletion causes aberrations in length and end-motif frequencies in plasma DNA\u003c/strong\u003e. \u003cem\u003eProc Natl Acad Sci U S A \u003c/em\u003e2019, \u003cstrong\u003e116\u003c/strong\u003e(2):641-649.\u003c/li\u003e\n\u003cli\u003eZhao Y, Wang J, Liang F, Liu Y, Wang Q, Zhang H, Jiang M, Zhang Z, Zhao W, Bao Y\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eNucMap: a database of genome-wide nucleosome positioning map across species\u003c/strong\u003e. \u003cem\u003eNucleic Acids Res \u003c/em\u003e2019, \u003cstrong\u003e47\u003c/strong\u003e(D1):D163-D169.\u003c/li\u003e\n\u003cli\u003eWei GH, Badis G, Berger MF, Kivioja T, Palin K, Enge M, Bonke M, Jolma A, Varjosalo M, Gehrke AR\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGenome-wide analysis of ETS-family DNA-binding in vitro and in vivo\u003c/strong\u003e. \u003cem\u003eThe EMBO journal \u003c/em\u003e2010, \u003cstrong\u003e29\u003c/strong\u003e(13):2147-2160.\u003c/li\u003e\n\u003cli\u003eDittmer J: \u003cstrong\u003eThe role of the transcription factor Ets1 in carcinoma\u003c/strong\u003e. \u003cem\u003eSeminars in cancer biology \u003c/em\u003e2015, \u003cstrong\u003e35\u003c/strong\u003e:20-38.\u003c/li\u003e\n\u003cli\u003eSaeki H, Kuwano H, Kawaguchi H, Ohno S, Sugimachi K: \u003cstrong\u003eExpression of ets-1 transcription factor is correlated with penetrating tumor progression in patients with squamous cell carcinoma of the esophagus\u003c/strong\u003e. \u003cem\u003eCancer \u003c/em\u003e2000, \u003cstrong\u003e89\u003c/strong\u003e(8):1670-1676.\u003c/li\u003e\n\u003cli\u003eSaeki H, Oda S, Kawaguchi H, Ohno S, Kuwano H, Maehara Y, Sugimachi K: \u003cstrong\u003eConcurrent overexpression of Ets-1 and c-Met correlates with a phenotype of high cellular motility in human esophageal cancer\u003c/strong\u003e. \u003cem\u003eInt J Cancer \u003c/em\u003e2002, \u003cstrong\u003e98\u003c/strong\u003e(1).\u003c/li\u003e\n\u003cli\u003eBaltrunaite K, Craig MP, Palencia Desai S, Chaturvedi P, Pandey RN, Hegde RS, Sumanas S: \u003cstrong\u003eETS transcription factors Etv2 and Fli1b are required for tumor angiogenesis\u003c/strong\u003e. \u003cem\u003eAngiogenesis \u003c/em\u003e2017, \u003cstrong\u003e20\u003c/strong\u003e(3):307-323.\u003c/li\u003e\n\u003cli\u003eAdamo P, Ladomery MR: \u003cstrong\u003eThe oncogene ERG: a key factor in prostate cancer\u003c/strong\u003e. \u003cem\u003eOncogene \u003c/em\u003e2016, \u003cstrong\u003e35\u003c/strong\u003e(4):403-414.\u003c/li\u003e\n\u003cli\u003eQiao G, Zhuang W, Dong B, Li C, Xu J, Wang G, Xie L, Zhou Z, Tian D, Chen G\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eDiscovery and validation of methylation signatures in circulating cell-free DNA for early detection of esophageal cancer: a case-control study\u003c/strong\u003e. \u003cem\u003eBMC medicine \u003c/em\u003e2021, \u003cstrong\u003e19\u003c/strong\u003e(1):243.\u003c/li\u003e\n\u003cli\u003eGilat N, Tabachnik T, Shwartz A, Shahal T, Torchinsky D, Michaeli Y, Nifker G, Zirkin S, Ebenstein Y: \u003cstrong\u003eSingle-molecule quantification of 5-hydroxymethylcytosine for diagnosis of blood and colon cancers\u003c/strong\u003e. \u003cem\u003eClinical epigenetics \u003c/em\u003e2017, \u003cstrong\u003e9\u003c/strong\u003e:70.\u003c/li\u003e\n\u003cli\u003eSong C-X, Yin S, Ma L, Wheeler A, Chen Y, Zhang Y, Liu B, Xiong J, Zhang W, Hu J\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages\u003c/strong\u003e. \u003cem\u003eCell research \u003c/em\u003e2017, \u003cstrong\u003e27\u003c/strong\u003e(10):1231-1242.\u003c/li\u003e\n\u003cli\u003eLi W, Zhang X, Lu X, You L, Song Y, Luo Z, Zhang J, Nie J, Zheng W, Xu D\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e5-Hydroxymethylcytosine signatures in circulating cell-free DNA as diagnostic biomarkers for human cancers\u003c/strong\u003e. \u003cem\u003eCell research \u003c/em\u003e2017, \u003cstrong\u003e27\u003c/strong\u003e(10):1243-1257.\u003c/li\u003e\n\u003cli\u003eShi X, Yu Y, Luo M, Zhang Z, Shi S, Feng X, Chen Z, He J: \u003cstrong\u003eLoss of 5-Hydroxymethylcytosine Is an Independent Unfavorable Prognostic Factor for Esophageal Squamous Cell Carcinoma\u003c/strong\u003e. \u003cem\u003ePloS one \u003c/em\u003e2016, \u003cstrong\u003e11\u003c/strong\u003e(4):e0153100.\u003c/li\u003e\n\u003cli\u003eMurata A, Baba Y, Ishimoto T, Miyake K, Kosumi K, Harada K, Kurashige J, Iwagami S, Sakamoto Y, Miyamoto Y\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eTET family proteins and 5-hydroxymethylcytosine in esophageal squamous cell carcinoma\u003c/strong\u003e. \u003cem\u003eOncotarget \u003c/em\u003e2015, \u003cstrong\u003e6\u003c/strong\u003e(27):23372-23382.\u003c/li\u003e\n\u003cli\u003eLi D, Zhang L, Liu Y, Sun H, Onwuka JU, Zhao Z, Tian W, Xu J, Zhao Y, Xu H: \u003cstrong\u003eSpecific DNA methylation markers in the diagnosis and prognosis of esophageal cancer\u003c/strong\u003e. \u003cem\u003eAging (Albany NY) \u003c/em\u003e2019, \u003cstrong\u003e11\u003c/strong\u003e(23):11640-11658.\u003c/li\u003e\n\u003cli\u003eSu J, Wu G, Ye Y, Zhang J, Zeng L, Huang X, Zheng Y, Bai R, Zhuang L, Li M\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eNSUN2-mediated RNA 5-methylcytosine promotes esophageal squamous cell carcinoma progression via LIN28B-dependent GRB2 mRNA stabilization\u003c/strong\u003e. \u003cem\u003eOncogene \u003c/em\u003e2021, \u003cstrong\u003e40\u003c/strong\u003e(39):5814-5828.\u003c/li\u003e\n\u003cli\u003eKit OI, Vodolazhskiy DI, Kolesnikov EN, Timoshkina NN: \u003cstrong\u003e[Epigenetic markers of esophageal cancer: DNA methylation]\u003c/strong\u003e. \u003cem\u003eBiomed Khim \u003c/em\u003e2016, \u003cstrong\u003e62\u003c/strong\u003e(5):520-526.\u003c/li\u003e\n\u003cli\u003eRice TW, Ishwaran H, Ferguson MK, Blackstone EH, Goldstraw P: \u003cstrong\u003eCancer of the Esophagus and Esophagogastric Junction: An Eighth Edition Staging Primer\u003c/strong\u003e. \u003cem\u003eJ Thorac Oncol \u003c/em\u003e2017, \u003cstrong\u003e12\u003c/strong\u003e(1):36-42.\u003c/li\u003e\n\u003cli\u003eLindgreen S: \u003cstrong\u003eAdapterRemoval: easy cleaning of next-generation sequencing reads\u003c/strong\u003e. \u003cem\u003eBMC Res Notes \u003c/em\u003e2012, \u003cstrong\u003e5\u003c/strong\u003e:337.\u003c/li\u003e\n\u003cli\u003eLangmead B, Salzberg SL: \u003cstrong\u003eFast gapped-read alignment with Bowtie 2\u003c/strong\u003e. \u003cem\u003eNat Methods \u003c/em\u003e2012, \u003cstrong\u003e9\u003c/strong\u003e(4):357-359.\u003c/li\u003e\n\u003cli\u003eEtherington GJ, Ramirez-Gonzalez RH, MacLean D: \u003cstrong\u003ebio-samtools 2: a package for analysis and visualization of sequence and alignment data with SAMtools in Ruby\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2015, \u003cstrong\u003e31\u003c/strong\u003e(15):2565-2567.\u003c/li\u003e\n\u003cli\u003eGrytten I, Rand KD, Nederbragt AJ, Storvik GO, Glad IK, Sandve GK: \u003cstrong\u003eGraph Peak Caller: Calling ChIP-seq peaks on graph-based reference genomes\u003c/strong\u003e. \u003cem\u003ePLoS Comput Biol \u003c/em\u003e2019, \u003cstrong\u003e15\u003c/strong\u003e(2):e1006731.\u003c/li\u003e\n\u003cli\u003eCavalcante RG, Sartor MA: \u003cstrong\u003eannotatr: genomic regions in context\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2017, \u003cstrong\u003e33\u003c/strong\u003e(15):2381-2383.\u003c/li\u003e\n\u003cli\u003eThorvaldsd\u0026oacute;ttir H, Robinson JT, Mesirov JP: \u003cstrong\u003eIntegrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration\u003c/strong\u003e. \u003cem\u003eBrief Bioinform \u003c/em\u003e2013, \u003cstrong\u003e14\u003c/strong\u003e(2):178-192.\u003c/li\u003e\n\u003cli\u003eRobinson JT, Thorvaldsd\u0026oacute;ttir H, Winckler W, Guttman M, Lander ES, Getz G, Mesirov JP: \u003cstrong\u003eIntegrative genomics viewer\u003c/strong\u003e. \u003cem\u003eNat Biotechnol \u003c/em\u003e2011, \u003cstrong\u003e29\u003c/strong\u003e(1):24-26.\u003c/li\u003e\n\u003cli\u003eShen L, Shao N, Liu X, Nestler E: \u003cstrong\u003engs.plot: Quick mining and visualization of next-generation sequencing data by integrating genomic databases\u003c/strong\u003e. \u003cem\u003eBMC Genomics \u003c/em\u003e2014, \u003cstrong\u003e15\u003c/strong\u003e:284.\u003c/li\u003e\n\u003cli\u003eLove MI, Huber W, Anders S: \u003cstrong\u003eModerated estimation of fold change and dispersion for RNA-seq data with DESeq2\u003c/strong\u003e. \u003cem\u003eGenome Biol \u003c/em\u003e2014, \u003cstrong\u003e15\u003c/strong\u003e(12):550.\u003c/li\u003e\n\u003cli\u003eTripathi S, Pohl MO, Zhou Y, Rodriguez-Frandsen A, Wang G, Stein DA, Moulton HM, DeJesus P, Che J, Mulder LCF\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eMeta- and Orthogonal Integration of Influenza \"OMICs\" Data Defines a Role for UBR4 in Virus Budding\u003c/strong\u003e. \u003cem\u003eCell Host Microbe \u003c/em\u003e2015, \u003cstrong\u003e18\u003c/strong\u003e(6):723-735.\u003c/li\u003e\n\u003cli\u003eGalili T, O'Callaghan A, Sidi J, Sievert C: \u003cstrong\u003eheatmaply: an R package for creating interactive cluster heatmaps for online publishing\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2018, \u003cstrong\u003e34\u003c/strong\u003e(9):1600-1602.\u003c/li\u003e\n\u003cli\u003eLiao Y, Smyth GK, Shi W: \u003cstrong\u003efeatureCounts: an efficient general purpose program for assigning sequence reads to genomic features\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2014, \u003cstrong\u003e30\u003c/strong\u003e(7):923-930.\u003c/li\u003e\n\u003cli\u003eSing T, Sander O, Beerenwinkel N, Lengauer T: \u003cstrong\u003eROCR: visualizing classifier performance in R\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2005, \u003cstrong\u003e21\u003c/strong\u003e(20):3940-3941.\u003c/li\u003e\n\u003cli\u003eRobin X, Turck N, Hainard A, Tiberti N, Lisacek F, Sanchez J-C, M\u0026uuml;ller M: \u003cstrong\u003epROC: an open-source package for R and S+ to analyze and compare ROC curves\u003c/strong\u003e. \u003cem\u003eBMC Bioinformatics \u003c/em\u003e2011, \u003cstrong\u003e12\u003c/strong\u003e:77.\u003c/li\u003e\n\u003cli\u003eWang Y, Song F, Zhu J, Zhang S, Yang Y, Chen T, Tang B, Dong L, Ding N, Zhang Q\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGSA: Genome Sequence Archive\u0026lt;sup/\u0026gt;\u003c/strong\u003e. \u003cem\u003eGenomics, proteomics \u0026amp; bioinformatics \u003c/em\u003e2017, \u003cstrong\u003e15\u003c/strong\u003e(1):14-18.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"esophageal cancer, early diagnosis, 5-hydroxymethylcytosine, low-passwhole-genome sequencing, cell-free DNA","lastPublishedDoi":"10.21203/rs.3.rs-1375061/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1375061/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eEsophageal cancer is one of globally high incidence and mortality disease. Its early stage has rarely obvious symptoms. Compared to conventional endoscopy diagnosis, liquid biopsy is an emerging non-invasive method for cancer early detection.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe enrolled 100 esophageal squamous cell carcinoma (ESCC) patients and 71 healthy individuals as a Southern China cohort and performed 5-hydroxymethylcytosine (5hmC) sequencing on their plasma cell-free DNA (cfDNA). A Northern cohort of cfDNA 5hmC dataset with 150 ESCC patients and 183 health individuals were downloaded for validation. A diagnostic model was firstly developed based on cfDNA 5hmC signatures and then improved by low-pass whole genome sequencing (WGS) features of cfDNA.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eConserved cfDNA 5hmC modification motifs were observed in the two independent ESCC cohorts. A diagnostic model with 273 5hmC features based on randomly-selected two-thirds samples of the Southern China cohort was validated independently in the left one-thirds samples and the whole Northern China cohort, achieved an AUC of 0.810 and 0.862 with sensitivities of 69.3\u0026ndash;74.3% and specificities of 82.4\u0026ndash;90.7%, respectively. The performance was well maintained in Stage I to Stage IV, with accuracy of 70%-100%, but low in Stage 0, with accuracy of 33.3%. Low-pass WGS of cfDNA improved the AUC to 0.934 with a sensitivity of 82.4%, a specificity of 88.2% and an accuracy of 84.3%, particularly significantly in Stage 0 with the accuracy up to 80%.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThis study suggests that the blood-based 5hmC integrated with low-pass WGS model could improve the accurate diagnosis of early stage ESCC, particularly for very early stage ESCC.\u003c/p\u003e","manuscriptTitle":"Plasma Cell-free DNA 5-Hydroxymethylcytosine and Whole-Genome Sequencing Signatures for Early Detection of Esophageal Cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-03-02 17:46:34","doi":"10.21203/rs.3.rs-1375061/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c5af3771-93bc-4ab2-acdf-c6cb485e3dde","owner":[],"postedDate":"March 2nd, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2022-03-09T20:29:13+00:00","versionOfRecord":[],"versionCreatedAt":"2022-03-02 17:46:34","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1375061","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1375061","identity":"rs-1375061","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0