Nanopore Direct RNA-Seq Reveals Widespread and Predictable Non-Coding Transcriptional Variations in DNA Methylation-deficient Arabidopsis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Nanopore Direct RNA-Seq Reveals Widespread and Predictable Non-Coding Transcriptional Variations in DNA Methylation-deficient Arabidopsis Ting Zhao This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6399991/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 31 Dec, 2025 Read the published version in Communications Biology → Version 1 posted You are reading this latest preprint version Abstract Long non-coding RNAs are regulatory RNAs with significantly high expression diversity within populations. In both plant and animal genomes, expression of long intergenic non-coding RNAs (lincRNAs) is associated with reduced genomic DNA methylation in non-coding regions. Whether such newly-activated lincRNAs are widespread in natural populations and have predictable patterns associated with natural selection has not been investigated. Here, we employed Oxford Nanopore Technology Direct RNA and DNA sequencing (ONT DRS and DDS) in DNA methylation-deficient Arabidopsis mutants ddm1 and met1 to generate 41 million high-quality long RNA reads. In total, 340 lincRNAs were found to be activated under defective DNA methylation, while 209 were constitutively expressed in both mutants and WT. Expression of ddm1 -activated lincRNAs was negatively correlated with DNA methylation levels. DNA-hypomethylation-activated lincRNAs can be detected in natural populations at low frequency and with high expression diversity. The ddm1 -activated and non- ddm1 -activated lincRNAs can be distinguished by Random Forest algorithm and Deep convolutional neural network with up to 70% and 91% accuracy. The integrated results suggest that dynamics of DNA methylation in non-coding regions are associated with non-coding RNAs in a predictable fashion with potential new function. This work advances a framework for leveraging epigenomic signatures in non-coding RNA discovery. Biological sciences/Genetics/Epigenetics/DNA methylation Biological sciences/Genetics/Genomics/Transcriptomics Direct RNA-sequencing non-coding RNA ddm1 met1 Arabidopsis hypomethylation machine learning Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Non-coding regions of genomes are reported to have great potential for transcription, and therefore play important roles in genomic evolution and regulation 1 , 2 . Long intergenic non-coding RNAs (lincRNAs), which are transcribed from intergenic regions and encode no recognizable proteins, are widespread in eukaryotic transcriptomes and exert a wide spectrum of influences within an organism 3 – 5 . Indeed, mounting evidence suggests that in plants, lincRNAs play regulatory roles in developmental and physiological processes including vernalization 6 , flowering development 7 , stress response 8 , 9 , and crop domestication 10 . LincRNAs as an unclassified pool are far more diverse than protein-coding genes (PCGs) in terms of expression activity, increased tissue specificity, and variability in expression across individuals 11 , 12 . Transcription of lincRNAs can result in novel genes and provide a route for expanding genome functionality without changing the number of genes 2 , 13 . From an evolutionary perspective, the flexibility of lincRNAs is important for generating genomic and phenotypic plasticity, such as constructive neutral evolution. It has been widely proposed that epigenetic variations can affect gene expression and cause visible phenotypic differences, allowing rapid adaptation to dynamic environmental changes 12 , 14 – 16 . Epigenome-wide association studies (EWAS) in natural populations along with studies using epigenetic recombinant inbred lines (epiRILs) have identified a large number of functional epigenetic variations 12 , 15 , 17 . More importantly, populational genetic analysis and comparative genome analysis have shown that a proportion of lincRNAs are under positive selection pressure 10 , 18 , 19 . That lincRNAs serve as a source of variation underlines the high potential of epigenome exploitation for breeding purposes. A relationship between lincRNAs and DNA methylation has been widely reported in animals and plants at both individual and population levels 20 . For example, population studies in the model plant Arabidopsis and in cancer cells have reported that lincRNA activation is associated with hypomethylation 19 , 21 . In plants, reprogramming of DNA methylation in the course of hybridization and polyploidization are reported to lead to accumulation of non-coding RNAs 22 – 24 . Studies of DNA methylation-related mutants have also characterized a large number of lincRNAs to be activated in multiple plant and animal species, including Oryza sativa , Arabidopsis thaliana , Zea mays , Brassica rapa , and Mus musculus 25 , 26 . Taken together, these studies indicate that genomes have a capability to integrate epigenetic modifications and transcription machinery on non-coding regions to create novel transcripts as a candidate pool for adaptive evolution within a short time. It is thus of great interest to identify and validate at genomic scale the flexibility of those non-coding regions having transcription potential. In Arabidopsis , DNA methylation mutants exhibiting different levels of DNA methylation deficiency have been well characterized 27 . Nucleosome remodelers of the DNA methylation 1 ( DDM1 )/Lsh family, which encode a protein similar to the chromatin-remodeling factor SWI2/SNF2, are required for maintenance of CG methylation (mCG) that is conserved in plants and animals 28 . Meanwhile, the protein encoded by methyltransferase 1 ( MET1 ) recognizes hemimethylated CG following DNA replication and methylates the naked cytosine in the daughter strand 29 . Subsequent maintenance of DNA methylation in heterochromatin requires DDM1 30 . Accordingly, ddm1 and met1 mutations have been shown to result in drastically decreased DNA methylation spreading on transposon elements (TEs) 31 – 33 . In addition, the hypomethylated status characteristic of ddm1 can be inherited even after eight generations 34 . These features make Arabidopsis ddm1 and met1 mutants excellent models for the investigation of lincRNA behaviors under a precondition of DNA hypomethylation 35 (Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020). About 80% of lincRNAs have 5′ caps and 3′ poly(A) tails, making them amenable for discovery by oligo (dT)-based RNA sequencing 36 . However, lincRNA sequences are also enriched with transposable element (TE) fragments 37 , 38 , which highly repetitive sequences cannot be accurately assembled by the next-generation sequencing (NGS, i.e., Illumina) technologies on which current studies rely; such technologies lack the necessary long-range information 39 , 40 . Direct RNA sequencing (DRS) offered by Oxford Nanopore Technologies (ONT) can sequence polyadenylated RNAs directly, without the recoding and amplification biases inherent to other sequencing methodologies 41 . Moreover, DRS can produce long reads (up to 30 kb, vs up to 600 bp at present for Illumina RNA-seq), which can provide for accurate transcript assemblies 42 . In addition, this method can detect the RNA modifications of m 6 A and m 5 C simultaneously, which introduce further epi-modification features in situ at single-molecule resolution 41 , 43 . Therefore, ONT DRS is an advanced technology suitable for the detailed study of lincRNAs. To investigate the underrepresented complexity of non-coding elements in the transcriptome, we leveraged the power of ONT DRS to characterize lincRNAs in Arabidopsis lines having null mutations of the DNA methylation regulators ddm1 and met1 . This approach was supplemented with short reads produced by Illumina RNA-seq. A large number of ddm1 -activated lincRNAs were identified and their occurrence in a natural population dissected. In addition, random forest classification algorithms were applied to directly learn DNA sequences traits from lincRNAs for the prediction of lincRNA sensitivity to activation. The presented models can serve as practical screening tools for mining causative non-coding variations in natural and artificial populations. Methods Study design Leaves from Arabidopsis thaliana Columbia (Col-0) wild type (WT) and strains respectively having mutations in DDM1 (TAIR Germplasm name: ddm1-2 , AT5G66750) and MET1 (TAIR Germplasm name: CS16394, met1-3 AT5G49160) ( ddm1 and met1 ) were prepared and sequenced using RNA Direct Sequencing, ONT DNA sequencing, and Illumina Strand-Specific RNA sequencing (ssRNA-seq). Plant materials and growth conditions Arabidopsis thaliana Col-0 WT, ddm1 , and met1 seeds were surface sterilized with 20% sodium hypochlorite (NaClO) for 10 min, followed by 70% ethanol for 5 min. Seeds were then washed four times in sterile distilled water, plated on growth medium, and placed in a growth chamber with 16 h light/8 h dark for 30 days. Thirty-day leaves (2 g) were harvested with four biological replicates and immediately immersed in liquid nitrogen. Two samples were used for an ONT DRS pilot run with a single flow cell (R9.4.1), while the remaining samples were used for ssRNA-seq. Nanopore direct RNA sequencing library preparation Total RNA from Arabidopsis leaves was isolated with Trizol according to the manufacturer’s protocol, and then polyA-selected using oligodT Dynabeads (Invitrogen). Approximately 200–500 ng of poly(A)-tailed mRNA was purified using a Dynabeads mRNA Purification Kit, followed by library construction using a Direct RNA Sequencing Kit. The libraries were sequenced on MinION MkIb instruments with a R9.4 sequencing microarray over 48–72 h using PromethION (Oxford Nanopore Technologies Ltd.). Each biological replicate and condition were sequenced independently in a different flow cell. Sequencing was performed by Wuhan Benagen Tech Solutions Company Limited, and the recommended scripts in MinKNOW were used to generate multi-fast5 read files with live base calling enabled. Alignment of ONT DRS Multi-fast5 reads were base-called using Guppy (v 0.1.10) with the default parameters 44 , which can achieve a calling accuracy of more than 99% 45 . LRs were separated into pass and fail folders; only reads in the pass folder (Phred > 7) were used in subsequent analysis. For each sample, ONT LRs were aligned to the Arabidopsis thaliana TAIR10 reference genome using Minimap2 (v 2.18, github.com/lh3/minimap2)), a long read local splice-aware aligner, with the following parameters: -ax splice -k14 -uf --secondary = no 46 . The resulting alignment files were converted to BAM format, sorted, and indexed by sequence coordinate with SAMtools (v1.6) 47 . Aligned reads with a secondary or supplementary aligned flag were discarded. Only reads with alignment flags 0 and 16 (primary mapping, forward strand and reverse strand of genome) were retained for further analysis. ONT reads were also aligned to the transcriptome (combined cDNA and non-coding RNA reference FASTA files) using Minimap2 (v 2.1)with the following parameters: -ax map-ont 46 . Generation of lincRNA annotations using ONT DRS For RNA annotation, we combined primary genome alignments from all libraries in each ONT data set, then assembled transcripts using StringTie2 (v2.0) in long read mode with the parameters -L 48 , 49 .The identified transcripts from each data set were compared with the set of annotated transcripts using StringTie2 (V2.0) 50 . For better annotation of isoforms, reads with similar splice junctions were grouped to generate a raw isoform annotation using FLAMES 51 . FLAMES can correct potential splice site and transcript start/end errors by comparing the raw isoform annotation against the reference annotation. Transcripts that have splice junctions (< 5 bp) and transcript start/end (< 100 bp) similar to the reference transcript were merged with the reference; thus, incomplete transcripts were collapsed. High-confidence isoforms were those identified with at least five supporting reads. Finally, we created a GTF file of the transcripts ( https://github.com/ZhaoTingzju/dRNA-seq ). Novel intergenic transcripts annotated with class code “u” by cuffcompare 52 were considered potential lincRNAs. The sequence of each transcript was extracted from its assembled form, and coding potential was assessed with CPC2 (v2.0b) 53 and Pfam_scan using HMMER (v 3.2.1, http://www.hmmer.org/ ). The obtained LR-RNA annotation, including coding genes, non-coding RNAs, and both novel and known annotated transcripts, are available online ( https://github.com/ZhaoTingzju/dRNA-seq ). Gene and transcript abundance estimation LR were aligned to the LR-RNA annotation and expression at gene level was quantified using StringTie2 (v2.0) 48 . To quantify genes, mapped reads were assigned to individual genes and counted by FeatureCount (v1.6.2) 54 with the arguments LongRead = TRUE (as recommended in the program help page when dealing with Nanopore data) and primaryOnly = TRUE to count primary alignments only. Differential gene expression analysis was conducted using the DESeq2 package 55 with the negative binomial statistic approach. The cutoff criteria for the selection of differentially expressed genes (DEGs) were: | log2-fold change | > 1 and false discovery rate (FDR) < 0.05. Poly(A) tail length estimation and alternative polyadenylation using ONT DRS ONT DRS sequencing requires a poly(A) tail for first adapter ligation. Poly(A) tail length was estimated from raw signal for each read using the poly(A) estimation function of the program Nanopolish (v0.11.0) 56 with “polya-merged” branch. Length estimates were only considered if the QC tag was Pass 56 , and were subsequently grouped by gene and isoform using an in-house Perl script. Detection of methylated adenosine and cytosine in ONT DRS Base-called multi-fast5 reads were converted to single-read fast5s using the Oxford Nanopore Technologies API ont_fasta5 ( https://github.com/nanoporetech/ont_fast5_api ). RNA modifications were identified using the software Tombo (v1.5.1) 44 . First, raw signal squiggles were assigned to the reference sequence using resquiggle . Next, modified base detection was carried out using detect modifications. The models of “5mC” and “ de novo ” were implemented separately to detect possible modifications in each read. MINES was used to detect m 6 A with default parameters ( https://github.com/YeoLab/MINES ). ONT DDS library preparation High-quality genomic DNA was isolated from fresh leaves using the CTAB method, and DNA quality and concentration were confirmed using 0.75% agarose gel electrophoresis, a NanoDrop One spectrophotometer (Thermo Fisher Scientific), and a Qubit 3.0 Fluorometer (Life Technologies, Carlsbad, CA, USA). Libraries were prepared with the SQK-LSK109 ligation kit according to the standard protocol. Each purified library was loaded onto a primed R9.4 Spot-On Flow Cell and sequenced using a PromethION sequencer (Oxford Nanopore Technologies, Oxford, UK) with 48-h runs at Wuhan Benagen Tech Solutions Company Limited, Wuhan, China. Base calling analysis of raw data was performed using the Oxford Nanopore Guppy software (v 0.1.10) 44 . Structure variation calling based on ONT DDS Multi-fast5 reads were base-called using Guppy (v 0.1.10) with the default parameters 44 . Passing LR reads were aligned to the Arabidopsis thaliana TAIR10 reference genome with Minimap2 (v 2.1) 46 , and SVs were called with cuteSV (v1.0.12), a sensitive SV detection implementation for LRs, from processed bam files using the suggested parameters for ONT: --max_cluster_bias_INS 100, --diff_ratio_merging_INS 0.3, -max_cluster_bias_DEL 100, --diff_ratio_merging_DEL 0.3 57 . SV information was stored in VCF format 58 ( https://github.com/ZhaoTingzju/dRNA-seq ). Illumina short-read RNA sequencing library preparation Strand-specific RNA-seq libraries were constructed following the TruSeq RNA sample preparation protocol as previously described with modifications 22 . Total RNA was treated with DNase first and then depleted of rRNA using complementary DNA oligos (IDT) and RNase H (Invitrogen, Waltham, MA, USA). Afterwards, RNA sequencing was conducted on the Illumina Hiseq 2500 platform with 150-bp paired-end reads according to the manufacturer’s protocol. Sequencing was performed by Wuhan Benagen Tech Solutions Company Limited. Gene and transcript abundance estimation based on SRs For SR-RNA annotation, the strand-specific paired-end 150-bp reads were first quality trimmed using Fastp 59 , then aligned to the Arabidopsis thaliana TAIR10 genome using Hisat2 with the following parameter: --dta 50 . Transcripts were subsequently assembled using StringTie (v2.0) with the –dta flag 48 . Differential gene expression analysis was conducted using the DESeq2 package with the negative binomial statistic approach 55 . Cutoff criteria for the selection of differentially expressed genes (DEGs) were: | log2-fold change | > 1 and false discovery rate (FDR) < 0.05. Analysis of publicly available data Publicly available SR transcriptomes from 675 diverse Arabidopsis thaliana accessions from the 1001 Genomes Project were downloaded from Gene Expression Omnibus (GEO, accession id: GSE80744). The reads were quality trimmed using Fastp 59 , then aligned to the Arabidopsis thaliana TAIR10 genome using Hisat2 with the following parameter: --dta 50 . Transcripts were subsequently quantified using StringTie (v2.0) based on the LR-RNA annotation 48 . The resulting data matrix, which contained expression values (FPKM) for all samples, was deposited in GitHub ( https://github.com/ZhaoTingzju/dRNA-seq ). Corresponding methylomes were likewise retrieved from GEO (accession id: GSE54292). The Illumina BS-Seq data was quality-checked with Fastp 59 . The clean reads were then aligned to the corresponding reference, and DNA methylation levels were calculated using Bismark 60 . Prediction of ddm1 -activated lincRNAs We implemented six machine learning models including Random Forest, XGBoost, SVM Linear, SVM RBF, Nearest Neighbors and Naive Bayes Classifier by using the library scikit learn in Python 3.10.9. The optimal parameter combination was selected based on algorithm accuracy as determined by five-fold cross-validation (CV). The scripts are available at https://github.com/ZhaoTingzju/dRNA-seq . In this study, we generated three broad categories of features: sequence features, gene structure and expression features, and RNA modification features. The sequence properties included the ratios of single nucleotides and dinucleotides, sequence length. Also, we upload sequence data on CPC2 (v2.0b) 53 to get more information including predicted coding probability and other features. TE annotation file was downloaded from TAIR10 ( http://www.arabidopsis.org ) to obtain the feature of the TE number in each lncRNA. The convolutional neural network was also implemented by using the library scikit learn in Python 3.10.9. The optimal parameter combination was selected based on algorithm accuracy as determined by ten-fold cross-validation. DNA sequences of both active and non-active lincRNAs were used as training input. the numbers of positive ( ddm1 -activated lincRNAs) and negative (non- ddm1 -activated lincRNAs) labels in the dataset are set to be same. Sequences of length 100 bp was converted into a 2-D matrix of size 4 × L, with each column a one-hot vector encoding the presence of the four DNA nucleotides A, C, G, and T. Method performance was evaluated using the ROC. The scripts are available at https://github.com/ZhaoTingzju/dRNA-seq . Statistics Statistical analysis was performed using Python 3.10.9 and R 4.4.1 software. Results Workflow of DRS and lincRNA annotation for Arabidopsis thaliana To investigate the fine details of lincRNAs, including expression variations, RNA modifications, and alternative poly(A) tails (Fig. 1 A), ONT DRS and Direct DNA Sequencing were performed on DNA isolated from leaf samples of Arabidopsis thaliana Col-0 wild type (WT) and ddm1 and met1 mutants. Illumina Strand-Specific RNA sequencing (ssRNA-seq) with 150-bp short reads (SRs) carried out in parallel on the same samples was also conducted for comparsion (Table S1 ). The analysis workflow for ONT DRS is shown in Fig. S1 . Base-calling and adapter trimming of the ONT long reads (LRs) was performed using Guppy (v0.1.10) 44 and yielded 41 million passing reads (Phred quality score > 7) from the six samples (Table S1 ). Taking advantage of the unprecedented read length of ONT DRS LRs, we assembled them for de novo RNA annotation. The analysis workflow is shown in Fig. 1 B. LRs from all libraries were combined for transcript assembling using the LR model of StringTie2 (v2.0) 48 , 49 . Transcripts with similar splice junctions (< 5 bp by default) and transcript start/end positions (< 100 bp) were further merged using FLAMES ( https://github.com/LuyiTian/FLAMES ), and those having at least five mapped LRs were retained. Assessment of DRS profiles and lincRNA annotation result The analysis of ONT DRS produced abundant high-quality reads. The median length of passed reads is 893 bp, with the maximum length being 14,248 bp. The read length distribution and read quality of LRs were similar to values reported in a previously published study of Arabidopsis thaliana 61 , but the total reads obtained was above 6 million (M) on average, which improved the total base coverage up to 7.4-fold more than the previous study, at about 0.8 M per library 61 . The LRs were mapped to the Arabidopsis thaliana reference genome TAIR10 ( http://www.arabidopsis.org ) and associated RNA annotations by Minimap2 (v2.1) 46 , yielding average mapping rates of 99.70% and 96.85%, respectively (Fig. S2 ; Table S2 ). The sequencing depth of each data set were further assessed with the frequencies of gene number identified and expression quantification using LRs. Although the total number of genes detected when using LRs was relatively less than the tally from SRs, at least 12,000 genes were efficiently evaluated by LRs (Fig. 2 A). LRs were randomly subsampled in amounts ranging from 0.5 to 5 M reads, and detections of expressed genes reached plateau phase at about 2.5 M reads (Fig. 2 B). Notably, read counts obtained from LR and SR sequencing were highly correlated, with Spearman’s correlation coefficient values ranging from 0.84 to 0.88 (Fig. 2 C; Fig. S3 ). Taken together, the profiles of the obtained ONT DRS data were consistent and allowed quantitative analysis of transcriptome. With regard to non-coding transcripts, less than 1% of LRs in the WT mapped to non-coding regions, but the amounts of non-coding reads were six-fold higher in ddm1 and two-fold higher in met1 (Fig. 2 D). This trend is in agreement with a previously reported observation of non-coding RNA activation in DNA methylation mutants 25 , which suggests that the LR libraries contained reliable RNAs from non-coding regions. The obtained LR-RNA annotation comprised 42,305 transcripts representing 15,925 genes. LincRNAs were further annotated by calculating coding potential and evaluating the presence of domains using coding potential calculator2 (CPC2, v2.0b) 53 and HMMER (v 3.2.1, http://hmmer.org/ ). Transcripts found to have protein coding potential or TE domains were excluded. Ultimately, 1,663 transcripts from 549 transcription models were annotated as lincRNAs. Of those, 98% were also represented in SR libraries (Fig. 2 E), confirming that ONT DRS can effectively capture the complex body of native full-length lincRNAs. Active expression of lincRNAs in ddm1 and met1 To examine the impact of DNA methylation on lincRNA transcriptional activity, PCGs and lincRNAs were quantified using normalized fragments per kilobase per million mapped reads (FPKM). For both LR and SR libraries, lincRNA expression levels in ddm1 (LR median = 6.23; SR median = 2.64) and met1 (LR median = 0.87; SR median = 0.43) were obviously higher than those in the WT (LR median = 0; SR median = 0) (typically P ≤ 2 × 10 − 16 , Wilcoxon test) (Fig. 3 A- 3 C). In contrast, the deficiency of DNA methylation in ddm1 and met1 mutants affected only a small fraction of PCG transcription in both LR and SR libraries, although the variations also reached statistical significance (typically P ≤ 2 × 10 − 16 , Wilcoxon test) (Fig. 3 D and 3 E). Trends in active lincRNAs in the hypomethylated mutants were confirmed by calculating the number of differentially expressed genes (DEGs). Of the annotated lincRNAs, 61.57% (n = 338) and 22.22% (n = 122) were differentially up-regulated in ddm1 and met1 , respectively (adjusted P < 0.05). For PCGs, the corresponding ratios were 2.41% (n = 662) and 1.55% (n = 426) in ddm1 and met1 , respectively. The same DEG analysis workflow was also applied to the SR libraries, and the majority of lincRNA DEGs were found to be shared between LR and SR libraries. Among the up-regulated lincRNAs detected in LR libraries, 310 (91.72%, out of 338) and 116 (95.08%, out of 122) were also differentially expressed in SR libraries in ddm1 and met1 (Fig. 3 F and S3). A representative lincRNA, MSTRG.12166 , was visualized in both LR and SR libraries using the Integrative Genomics Viewer (IGV) (Fig. 3 G). This transcript was specifically expressed in ddm1 , but was not expressed in the corresponding WT or met1 samples. Similarly, Fig. 3 H illustrates a second representative lincRNA, MSTRG.12270 , that was actively expressed in ddm1 and met1 but not expressed in WT. All told, the ONT DRS profiling confirmed our previous conclusions regarding the impact of DNA hypomethylation on activation of lincRNAs in multiple organisms, which might alter transcriptome plasticity 25 . Thus, the lincRNAs were classified according to expression specificity, yielding 340 ddm1 -activated lincRNAs and 209 non- ddm1 -activated lincRNAs for further analysis (Table S3 ). Patterns of RNA modifications and poly(A) tails on lincRNAs The impacts of DNA hypomethylation on lincRNAs were further examined in relation to RNA m 6 A, m 5 C modification and poly(A) tail length. In total, an average of 37,808 high-confidence m 6 A loci were identified using Tombo (v 1.5.1) 44 and MINES 62 . Each transcript had on average four m 6 A sites, which is consistent with a previous result in Populus trichocarpa 63 . Meanwhile, an average of 1,954,236 m 5 C loci were detected using Tombo (v 1.5.1) 44 . In the ddm1 mutant, the ratios of transcripts with m 6 A modification to those with m 5 C was similar between PCGs (m 6 A 68.90%; m 5 C 67.64%) and ddm1 -activated lincRNAs (m 6 A 67.65%; m 5 C 63.83%), and slightly lower in non- ddm1 activated lincRNAs (m 6 A 36.85%; m 5 C 42,59%) (Fig. 4 A). Poly(A) tails are important for transcription and transcript stability 64 , 65 . Here we examined poly(A) tail length using the low-variance ionic current signal associated with the 3’ end in ONT DRS 56 , 66 . This method identified 15,014 transcripts with poly(A) tails (Fig. 4 A), which accounted for 95.10% of PCGs, 98.53% of ddm1 -activated lincRNAs, and 93.78% of non- ddm1 -activated lincRNAs. Overall, non-coding RNAs featured longer 3’ UTRs ( ddm1 -activated, mean 79 nt; non- ddm1 -activated, mean 82 nt) than did PCGs (mean 68 nt) ( P ≤ 2 × 10 − 16 , Wilcoxon test) (Fig. 4 B). The length of poly(A) tails was negatively correlated with transcriptional activity (Fig. S5 ). consistent with a previous report that poly(A) length negatively correlates with expression 67 . DNA methylation according to ONT DDS is negatively correlated with transcriptional activity at lincRNA loci To investigate the quantitative relationships between DNA methylation and lincRNA transcriptional activity, the ONT DDS technology was employed to quantify DNA methylation levels without chemical treatment and structure variation. LR DNA sequencing with the ONT platform was performed using the same plant materials. Base calling was carried out using Guppy (v 0.1.10) [32], and yielded 3.10 M reads (Phred score > 7) with an average length of 14 kb (Table S1 ). The longest read obtained was 208,904 bp. The mean coverage of each sample was 61× across the genome. The reads were mapped to the Arabidopsis thaliana genome TAIR10 by Minimap2 (v 2.1) [34], which revealed 189 structural variations in ddm1 and 118 in met1 mutants in compare with WT (Fig. S6 ). Structural variations were observed in only small fractions of ddm1 -activated lincRNAs (3.82%, n = 13) and of non- ddm1 -activated lincRNAs (2.87%, n = 6); accordingly, there was no statistical enrichment of structural variation in ddm1 -activated lincRNAs ( P = 0.36, Fisher's Exact Test). (Table S4 ). To examine the association of lincRNA activity with DNA methylation level, genome-wide DNA methylation was first confirmed in ddm1 and met1 ( p < 2.22 × 10 − 16 ) in compared with WT (Fig. 5 A) based on the calling of mC (methylated cytosine) frequency in reads by Nanopolish 44 . Comparison of gene body DNA methylation in ddm1 -activated lincRNAs between ddm1 (mean: 11.98%) and WT (mean: 79.17%) revealed a 67.20% decrease in the mutant (Fig. 5 B). Meanwhile, non- ddm1 -activated lincRNAs and PCGs exhibited respective decreases of only 16.97% (Fig. 5 C) and 5.58% (Fig. 5 D). These trends confirmed that the activation of lincRNAs in ddm1 was significantly correlated with alteration of DNA modifications. This is further illustrated by a representative ddm1 -activated lincRNA, MSTRG.12166 , at which locus the genomic sequences in WT and ddm1 were identical, but RNA activation was specific to the ddm1 mutant and corresponded with reduced DNA methylation (Fig. 5 E). Hence, CG methylation is most likely a major obstacle to the expression of lincRNAs. ddm1 -activated lincRNAs can be detected in Arabidopsis natural populations The Arabidopsis ddm1 mutant expresses a group of activated lincRNAs that have structural and epi-modification features similar to PCGs (Fig. 4 A); this finding indicates that non-coding regions of the genome might have the potential to generate new genes when released from DNA methylation. This hypothesis inspired us to examine the expression of ddm1 -activated lincRNAs in natural populations so as to confirm the deduced lincRNA behavioral patterns at large scale. To further broaden the biological scope of this analysis, we obtained from a public database the leaf transcriptome profiles of 728 Arabidopsis natural accessions from the 1001 Genomes Project, along with accompanying original information concerning geographic distribution 14 (Table S5 ). We then investigated the expression of ddm1 -activated and non- ddm1 -activated lincRNAs in those accessions, with expression defined as having FPKM > 1. Interestingly, the majority of annotated lincRNAs (68.31%, n = 375) were expressed in at least one of the tested accessions (Fig. 6 A). In addition, ddm1 -activated lincRNAs originally identified under condition of hypomethylation were found to be active in this natural population, although only at low frequency (Fig. 6 B). The expression variation of ddm1 -activated lincRNAs was significantly larger than that of either non- ddm1 -activated lincRNAs or PCGs, determined by calculating the average coefficient variation (CV, P ≤ 2 × 10 − 16 , Wilcoxon test) (Fig. 6 C). For example, one ddm1 -activated lincRNA with multiple exons, MSTRG.19779 , was not expressed in most accessions (FPKM < 1), similar to its state in Col-0 WT. However, it was highly expressed in at least three accessions, IP-San-10, Gd-1, Corig-1, similar to the situation in ddm1 , with a high level of expression and a stable RNA splicing pattern (Fig. 6 D). Moreover, the region showed significant decrease of DNA methylation in those accessions, including CG, CHG, CHH methylations (Fig. 6 E). These results suggest that in natural populations, many accessions also present activation of lincRNA expression similar to what occurred in the ddm1 mutant with genome-wide DNA hypomethylation. Prediction of ddm1 -activated lincRNAs by machine learning and deep learning As noted above, ddm1 -activated lincRNAs differ from non- ddm1 -activated lincRNAs in terms of expression levels (Fig. 3 C) as well as epigenetic landscapes (Fig. 4 A). These patterns provide a fundamental basis for classifying lincRNAs into distinct groups. We then hypothesized that the observed activation patterns of lincRNAs are predictable, and through sufficient information, they can be classified into ddm1 -activated or non- ddm1 -activated categories. Thus, thirty-five features per lincRNA were constructed and used in six machine learning classification models including Random Forest, XGBoost (eXtreme Gradient Boosting), SVM Linear (Support Vector Machine with Linear Kernel), SVM RBF (Support Vector Machine with Radial Basis Function Kernel), Nearest Neighbors and Naive Bayes Classifier (Table S6 ). Model performance was evaluated through five iterations of five-fold cross-validation and the result showed that Random Forest, XGBoost and SVM Linear models reached the best three prediction accuracy (Fig. 7 A and S7). Among these models, Random Forest achieved the highest AUC-ROC of 0.83 and average precision of 0.88, followed by XGBoost (AUC-ROC = 0.81, AP = 0.86) and SVM Linear (AUC-ROC = 0.71, AP = 0.76) (Fig. 7 B and 7 C). In addition, we ranked all features within each model in order of importance and presented the top seven features (Fig. S7 D). Notably, the number of transposable elements (TE) consistently ranked among the top two in the best three models (Random Forest, XGBoost and SVM Linear), indicating a potential relationship between TE abundance and lincRNA activation (Fig. S7 D). Inspired by the successful application of deep learning in direct DNA sequence, a deep convolutional neural network (DCNN) was then implemented for classifying lincRNAs as ddm1 -activated or not based on the 100 bp sequence context (Fig. 7 D). The performance of this fully sequence-based model was evaluated using a 10 times five-fold cross-validation procedure, in which it achieved an area under receiver operating characteristic value of 0.91, and average precision 0.91 (Fig. 7 E and 7 F). The DCNN model demonstrated better accuracy in the classification of ddm1 -activated RNA compared to conventional machine learning models. This capability stems from DCNN's hierarchical architecture, where convolutional layers directly and automatically extracted multi-scale features related to lncRNA transcriptional activation from DNA sequences. Conclusion In this study, we demonstrated that ONT DRS technology is efficient for identifying at high confidence lincRNA transcript structures that have been overlooked by short-read sequencing. LincRNAs were finely mapped and epi-modifications examined at single molecule resolution. The ddm1 -activated lincRNAs can also be activated in natural populations. Using thirty-five features, six machine learning models were constructed among which Random Forest model enabled prediction of transcript activation in the absence of DNA methylation with 70% accuracy. All told, with the integration of lincRNA activation patterns in populations and their predictable sequence patterns identified with machine-learning technology, we propose the use of lincRNAs and associated DNA methylation patterns as new molecular markers in the development of crop breeding strategies. Discussion Advantages of long-read sequencing technology in investigating non-coding RNAs At present, ONT DRS is mainly applied in studies of protein coding genes to identify structure variations and epigenetic modifications based on its long-read features and the accessorized detection of RNA modifications. We demonstrate that ONT DRS can also be utilized to identify and quantify lincRNAs with high accuracy. It is known that lincRNAs are transcribed from TE-enriched intergenic regions 37 , which feature repetitive sequences and genomic SVs associated with TE mobilizations; this makes it challenging to use short-read technology to navigate the structure of lincRNAs on both RNA and DNA levels. The DRS capability of ONT has been widely adapted for studies on alternative splicing 39 , 61 . Native RNA molecules may carry a variety of base modifications, such as m 6 A and m 5 C 68 , which play roles in the stabilization of RNA transcription and translation. By taking advantage of ONT LR technology, such epi-modifications can be characterized without technique bias. Thus, application of Nanopore-based DRS technology can be expected to greatly advance our understanding of non-coding RNAs and their complex structure, polyadenylation, and modification. In humans, m 6 A modification has been reported for lincRNAs, MALAT1 , TUG1 , TPT1 , ACTB , and BSG 69 . In this study, we found that most lincRNAs in Arabidopsis featured m 6 A and m 5 C modification. In addition to the above, the long reads from DRS have clear advantages in distinguishing homologs within a complex genome featuring redundant sequences, such as in polyploid species and those sequences from repetitive regions enriched for TEs. ONT DNA sequencing offers exceptionally long reads that permit direct sequencing through regions of the genome previously deemed inaccessible. Although the data presented here are preliminary, we still reveal a clear and systematic difference between groups of molecules at high confidence, namely ddm1 -activated lincRNAs and non- ddm1 -activated lincRNAs. De novo expression of lincRNAs found in natural populations The activation of lincRNAs seen in ddm1 and met1 mutants can also be observed in natural populations, supporting that the non-hypomethylated genome retains the capability to employ its non-coding regions and transcribe RNA from them under given conditions. In the last decade, substantial efforts have been conducted to annotate lincRNAs in multiple species, it still remains a great challenge to identify the functional effects of lincRNA variants in genetic terms 70 . A fundamental barrier is the identification of regulatory DNA elements on lincRNA loci, which requires a classification system for lincRNA identification and correlation with expression diversity. Further classification of such lincRNAs will be facilitated by investigating lincRNA expression polymorphisms in natural populations and identifying de novo active lincRNAs that have not yet been acted on. Prediction of ddm1 -activated lincRNAs using machine learning and deep learning Machine learning algorithms have been applied to predict genes transcriptionally patterns by training on biochemical features. Here we successfully predict hypomethylation-responsive lincRNAs, which can exhibit more expression diversity. The rank of feature importance within each model helps to reveal possible causes between lincRNA activation and methylation. The dynamics of epigenetic modifications on DNA and RNA are proposed to occur in spontaneous response to environmental and developmental stimuli, in keeping with their roles as constituting the flexible part of the genome. Several DNA methylation loci introduced by TE-associated effects have been reported to serve as critical regulators in crops 71 . The integration of machine learning provides an appoach for epigenomics research on lincRNA. Expression of lincRNAs might occur as a side product of DNA hypomethylation, but hypomethylation-associated lincRNAs have also been reported from crop studies. For example, GhDAN1 in cotton is a non-coding gene reported to regulate drought response by forming a DNA: RNA triplex in leaves. More interestingly, GhDAN1 underwent neofunctionalization in allotetraploid cotton by being released from a hypomethylated chromatin region 23 . In addition, a non-coding RNA ( At2g0600 ) has been identified as a novel epiallele involved in flowering through simulating selection in a rapidly changing environment and likewise associated with DNA methylation 19 . Therefore, lincRNAs provide a kind of rapid heritable variation on which natural selection can act. Convolutional neural network models have been successfully applied to predict mRNA transcriptional activation 72 and DNA methylation levels 73 based on DNA sequences. Here we successfully predict hypomethylation-related lincRNAs, which can exhibit more expression diversity. Further classification of such lincRNAs will be facilitated by investigating lincRNA expression polymorphisms in natural populations and identifying de novo active lincRNAs that have not yet been acted on. Data availability All raw and processed sequencing data generated in this study have been submitted to the NCBI. The high-quality reads and RNA reads from ONT were deposited in the SRA database under accession number BioProject: PRJNA780475. Declarations Acknowledgements This work was financially supported in part by grants from the National Natural Science Foundation of China (NSFC 32000379). We thank Tianzhen Zhang from Zhejiang University for providing insightful suggestions on the populational analysis. Author information Authors and Affiliations Zhejiang Provincial Key Laboratory of Crop Genetic Resources, Institute of Crop Science, Plant Precision Breeding Academy, College of Agriculture and Biotechnology, Zhejiang University, Hangzhou, China Wanghong Shi, Mengke Zhang, Yupeng Hao, Ke Nie, Xueying Guan & Ting Zhao Hainan Institute of Zhejiang University, Building 11, Yonyou Industrial Park, Yazhou Bay Science and Technology City, Yazhou District, Sanya, Hainan, 572025, China Wanghong Shi, Luyao Wang, Mengke Zhang, Yupeng Hao & Ke Nie Dry land Farming Institute, Hebei Academy of Agricultural and Forestry Sciences, Hengshui, 053000, China Na Zhou Contributions T. Z. conceptualized the project. X. G., T. Z., L. W., Y. H. and K. N. performed the experiment. W. S., T. Z., N. Z., M. Z. and X. G. analyzed the data. W. S., T. Z. and X. G. prepared the manuscript. All authors read and approved the final manuscript. References Lee, H., Zhang, Z., Krause, H.M.: Long Noncoding RNAs and Repetitive Elements: Junk or Intimate Evolutionary Partners? Trends Genet. 35 , 892–902 (2019) Palazzo, A.F., Koonin, E.V.: Functional Long Non-coding RNAs Evolve from Junk Transcripts. Cell. 183 , 1151–1161 (2020) Waseem, M., Liu, Y., Xia, R.: Long Non-Coding RNAs, the Dark Matter: An Emerging Regulatory Component in Plants. Int. J. Mol. Sci. 22 (2020) Iyer, M.K., et al.: The landscape of long noncoding RNAs in the human transcriptome. Nat. Genet. 47 , 199–208 (2015) Andergassen, D., Rinn, J.L.: From genotype to phenotype: genetics of mammalian long non-coding RNAs in vivo. Nat. Rev. Genet. 23 , 229–243 (2022) Heo, J.B., Sung, S.: Vernalization-mediated epigenetic silencing by a long intronic noncoding RNA. Science. 331 , 76–79 (2011) Zhao, Y., et al.: Natural temperature fluctuations promote COOLAIR regulation of FLC. Genes Dev. 35 , 888–898 (2021) Tao, X., et al.: Neofunctionalization of a Polyploidization-Activated Cotton Long intergenic non-coding RNA DAN1 During Drought Stress Regulation. Plant Physiol (2021) Wang, D., et al.: Transposable elements (TEs) contribute to stress-related long intergenic noncoding RNAs in plants. Plant. J. 90 , 133–146 (2017) Zheng, X.M., et al.: Genome-wide analyses reveal the role of noncoding variation in complex traits during rice domestication. Sci. Adv. 5 , eaax3619 (2019) Mele, M., et al.: Human genomics. The human transcriptome across tissues and individuals. Science. 348 , 660–665 (2015) de Goede, O.M., et al.: Population-scale tissue transcriptomics maps long non-coding RNAs to complex disease. Cell. 184 , 2633–2648e19 (2021) Vakirlis, N., et al.: De novo emergence of adaptive membrane proteins from thymine-rich genomic sequences. Nat. Commun. 11 , 781 (2020) Kawakatsu, T., et al.: Epigenomic Diversity in a Global Collection of Arabidopsis thaliana Accessions. Cell. 166 , 492–505 (2016) Furci, L., et al.: Identification and characterisation of hypomethylated DNA loci controlling quantitative resistance in Arabidopsis. Elife 8 (2019) Kooke, R., et al.: Epigenetic basis of morphological variation and phenotypic plasticity in Arabidopsis thaliana. Plant. Cell. 27 , 337–348 (2015) He, L., et al.: A naturally occurring epiallele associates with leaf senescence and local climate adaptation in Arabidopsis accessions. Nat. Commun. 9 (2018) Glenfield, C., McLysaght, A.: Pseudogenes Provide Evolutionary Evidence for the Competitive Endogenous RNA Hypothesis. Mol. Biol. Evol. 35 , 2886–2899 (2018) Schmid, M.W., et al.: Contribution of epigenetic variation to adaptation in Arabidopsis. Nat. Commun. 9 , 4446 (2018) Zhao, T., et al.: Population-wide DNA methylation polymorphisms at single-nucleotide resolution in 207 cotton accessions reveal epigenomic contributions to complex traits. Cell. Res. 34 , 859–872 (2024) Wang, Z., et al.: lncRNA Epigenetic Landscape Analysis Identifies EPIC1 as an Oncogenic lncRNA that Interacts with MYC and Promotes Cell-Cycle Progression in Cancer. Cancer Cell. 33 , 706–720e9 (2018) Zhao, T., et al.: LncRNAs in polyploid cotton interspecific hybrids are derived from transposon neofunctionalization. Genome Biol. 19 , 195 (2018) Tao, X., et al.: Neofunctionalization of a polyploidization-activated cotton long intergenic non-coding RNA DAN1 during drought stress regulation. Plant. Physiol. 186 , 2152–2168 (2021) Wang, P., et al.: Global Profiling of lncRNAs Expression Responsive to Allopolyploidization in Cucumis. Genes (Basel) 11 (2020) Zhao, T., et al.: Absence of CG methylation alters the long noncoding transcriptome landscape in multiple species. FEBS Lett. 595 , 1734–1747 (2021) Tan, F., et al.: DDM1 Represses Noncoding RNA Expression and RNA-Directed DNA Methylation in Heterochromatin. Plant. Physiol. 177 , 1187–1197 (2018) Stroud, H., Greenberg, M.V., Feng, S., Bernatavichute, Y.V., Jacobsen, S.E.: Comprehensive analysis of silencing mutants reveals complex regulation of the Arabidopsis methylome. Cell 152, 352 – 64 (2013) Du, J.M., Johnson, L.M., Jacobsen, S.E., Patel, D.J.: DNA methylation pathways and their crosstalk with histone methylation. Nat. Rev. Mol. Cell Biol. 16 , 519–532 (2015) Law, J.A., Jacobsen, S.E.: Establishing, maintaining and modifying DNA methylation patterns in plants and animals. Nat. Rev. Genet. 11 , 204–220 (2010) Osakabe, A., et al.: The chromatin remodeler DDM1 prevents transposon mobility through deposition of histone variant H2A.W. Nat. Cell Biol. 23 , 391–U177 (2021) Vongs, A., Kakutani, T., Martienssen, R.A., Richards, E.J.: Arabidopsis thaliana DNA methylation mutants. Science. 260 , 1926–1928 (1993) Kankel, M.W., et al.: Arabidopsis MET1 cytosine methyltransferase mutants. Genetics. 163 , 1109–1122 (2003) Lippman, Z., et al.: Role of transposable elements in heterochromatin and epigenetic control. Nature. 430 , 471–476 (2004) Quadrana, L., Colot, V.: Plant Transgenerational Epigenetics. Annu. Rev. Genet. 50 , 467–491 (2016) Tao, P., et al.: The overexpression of lncRNA MEG3 inhibits cell viability and invasion and promotes apoptosis in ovarian cancer by sponging miR-205-5p. Int. J. Clin. Exp. Pathol. 13 , 869–879 (2020) Ulitsky, I., Bartel, D.P.: lincRNAs: Genomics, Evolution, and Mechanisms. Cell. 154 , 26–46 (2013) Kapusta, A., et al.: Transposable elements are major contributors to the origin, diversification, and regulation of vertebrate long noncoding RNAs. PLoS Genet. 9 , e1003470 (2013) Babarinde, I.A., et al.: Transposable element sequence fragments incorporated into coding and noncoding transcripts modulate the transcriptome of human pluripotent stem cells. Nucleic Acids Res. 49 , 9132–9153 (2021) Wang, X., et al.: Full-length transcriptome reconstruction reveals a large diversity of RNA and protein isoforms in rat hippocampus. Nat. Commun. 10 , 5009 (2019) Ebbert, M.T.W., et al.: Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight. Genome Biol. 20 , 97 (2019) Garalde, D.R., et al.: Highly parallel direct RNA sequencing on an array of nanopores. Nat. Methods. 15 , 201–206 (2018) Seki, M., et al.: Evaluation and application of RNA-Seq by MinION. DNA Res. 26 , 55–65 (2019) Qin, H., et al.: DENA: training an authentic neural network model using Nanopore sequencing data of Arabidopsis transcripts for detection and quantification of N(6)-methyladenosine on RNA. Genome Biol. 23 , 25 (2022) Wick, R.R., Judd, L.M., Holt, K.E.: Performance of neural network basecalling tools for Oxford Nanopore sequencing. Genome Biol. 20 , 129 (2019) Xie, S., et al.: Applications and potentials of nanopore sequencing in the (epi)genome and (epi)transcriptome era. Innov. (N Y). 2 , 100153 (2021) Li, H.: Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 34 , 3094–3100 (2018) Li, H., et al.: The Sequence Alignment/Map format and SAMtools. Bioinformatics. 25 , 2078–2079 (2009) Pertea, M., Kim, D., Pertea, G.M., Leek, J.T., Salzberg, S.L.: Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat. Protoc. 11 , 1650–1667 (2016) Kovaka, S., et al.: Transcriptome assembly from long-read RNA-seq alignments with StringTie2. Genome Biol. 20 (2019) Trapnell, C., et al.: Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat. Protoc. 7 , 562–578 (2012) Tian, L., et al.: (2020) Trapnell, C., et al.: Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat. Biotechnol. 28 , 511–515 (2010) Kang, Y.J., et al.: CPC2: a fast and accurate coding potential calculator based on sequence intrinsic features. Nucleic Acids Res. 45 , W12–W16 (2017) Liao, Y., Smyth, G.K., Shi, W.: featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30, 923 – 30 (2014) Love, M.I., Huber, W., Anders, S.: Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 15 , 550 (2014) Workman, R.E., et al.: Nanopore native RNA sequencing of a human poly(A) transcriptome. Nat. Methods. 16 , 1297–1305 (2019) Jiang, T., et al.: Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 21 , 189 (2020) Danecek, P., et al.: The variant call format and VCFtools. Bioinformatics. 27 , 2156–2158 (2011) Chen, S., Zhou, Y., Chen, Y., Gu, J.: fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 34 , i884–i890 (2018) Krueger, F., Andrews, S.R.: Bismark: a flexible aligner and methylation caller for Bisulfite-Seq applications. Bioinformatics. 27 , 1571–1572 (2011) Zhang, S., et al.: New insights into Arabidopsis transcriptome complexity revealed by direct sequencing of native RNAs. Nucleic Acids Res. 48 , 7700–7711 (2020) Lorenz, D.A., Sathe, S., Einstein, J.M., Yeo, G.W.: Direct RNA sequencing enables m(6)A detection in endogenous transcript isoforms at base-specific resolution. RNA. 26 , 19–28 (2020) Gao, Y., et al.: Quantitative profiling of N(6)-methyladenosine at single-base resolution in stem-differentiating xylem of Populus trichocarpa using Nanopore direct RNA sequencing. Genome Biol. 22 , 22 (2021) Lim, J., Lee, M., Son, A., Chang, H., Kim, V.N.: mTAIL-seq reveals dynamic poly(A) tail regulation in oocyte-to-embryo development. Genes Dev. 30 , 1671–1682 (2016) Lima, S.A., et al.: Short poly(A) tails are a conserved feature of highly expressed genes. Nat. Struct. Mol. Biol. 24 , 1057–1063 (2017) Loman, N.J., Quick, J., Simpson, J.T.: A complete bacterial genome assembled de novo using only nanopore sequencing data. Nat. Methods. 12 , 733–U51 (2015) Xiang, K., Bartel, D.P.: The molecular basis of coupling between poly(A)-tail length and translational efficiency. Elife 10 (2021) Ameur, A., Kloosterman, W.P., Hestand, M.S.: Single-Molecule Sequencing: Towards Clinical Applications. Trends Biotechnol. 37 , 72–85 (2019) Liu, N., et al.: Probing N6-methyladenosine RNA modification status at single nucleotide resolution in mRNA and long noncoding RNA. RNA. 19 , 1848–1856 (2013) Huttenhofer, A., Schattner, P., Polacek, N.: Non-coding RNAs: hope or hype? Trends Genet. 21 , 289–297 (2005) Xu, G., et al.: Evolutionary and functional genomics of DNA methylation in maize domestication and improvement. Nat. Commun. 11 , 5539 (2020) Washburn, J.D., et al.: Evolutionarily informed deep learning methods for predicting relative transcript abundance from DNA sequence. Proc. Natl. Acad. Sci. U S A. 116 , 5542–5549 (2019) Zeng, H., Gifford, D.K.: Predicting the impact of non-coding variants on DNA methylation. Nucleic Acids Res. 45 , e99 (2017) Additional Declarations There is NO Competing Interest. Supplementary Files FigS1.pdf Supplementary Figure 1 FigS2.pdf Supplementary Figure 2 FigS3.pdf Supplementary Figure 3 FigS4.pdf Supplementary Figure 4 FigS5.pdf Supplementary Figure 5 FigS6.pdf Supplementary Figure 6 FigS7.pdf Supplementary Figure 7 TableS.xlsx Supplementary Table DescriptionofSupplementaryFiles.docx The description of supplementary files Cite Share Download PDF Status: Published Journal Publication published 31 Dec, 2025 Read the published version in Communications Biology → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6399991","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":450818385,"identity":"dd33fd87-da1d-4338-8e25-d3ef652f063f","order_by":0,"name":"Ting Zhao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIiWNgGAWjYDACZiBmbGBg4GdgSIByCQAemBbJBqK1MEC1GByAW0oA2LPzHn75c4ednPH5A88kGCqsExvYzx4g4DC+NAvJM8nGZgcOpEkwnElPbODJSyCghcfMwLCNOXHbwYY0Cca2w4kNEjwGhLUkttUnbm5mAGr5R5wW4wcHgYZvYANpaSBGy2EeM8bGtuPGEmcYki0SjqUbt/Hk4NfC3n/G+OPPtmo5/v4ziTc+1FjL9rOfwa8FCNgkoBYmgCOTjZB6IGD+ALXwABGKR8EoGAWjYCQCAPlqPtk8T6NHAAAAAElFTkSuQmCC","orcid":"","institution":"Zhejiang University","correspondingAuthor":true,"prefix":"","firstName":"Ting","middleName":"","lastName":"Zhao","suffix":""}],"badges":[],"createdAt":"2025-04-08 06:46:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6399991/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6399991/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s42003-025-09454-1","type":"published","date":"2025-12-31T05:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":82206117,"identity":"bca9c4f2-8342-4c64-8a98-666ff8ccc93c","added_by":"auto","created_at":"2025-05-07 17:28:45","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":51733,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow of this study.\u003cstrong\u003e A:\u003c/strong\u003e Schematic graph showing the study workflow. Leaves of \u003cem\u003eArabidopsis thaliana\u003c/em\u003e Col-0 (WT), \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e mutants were harvested for direct native RNA and DNA sequencing using the ONT platform. RNA reads were assembled for annotation. Genomic features of m\u003csup\u003e6\u003c/sup\u003eA and m\u003csup\u003e5\u003c/sup\u003eC modification, poly(A) tail, and structural variations (SV) were collected for integrative analysis and machine learning work.\u003cstrong\u003e B: \u003c/strong\u003eWorkflow for lincRNA annotation based on ONT DRS.\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/9a1a117e6ae9d2de17b48706.jpg"},{"id":82206118,"identity":"f2b431d2-189b-4d30-9d2d-beca2ff87234","added_by":"auto","created_at":"2025-05-07 17:28:45","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":56676,"visible":true,"origin":"","legend":"\u003cp\u003eSequencing of \u003cem\u003eArabidopsis\u003c/em\u003e \u003cem\u003eddm1\u003c/em\u003e, \u003cem\u003emet1\u003c/em\u003e, and WT genotypes using ONT DRS technology.\u003cstrong\u003e A: \u003c/strong\u003eSummary of ONT sequencing data volume.\u003cstrong\u003e B: \u003c/strong\u003eSequencing saturation analysis of ONT DRS mapping in each genotype. Y-axis represents the number of genes covered at a given number of reads (X-axis).\u003cstrong\u003e C: \u003c/strong\u003eSpearman’s \u003cem\u003eR\u003c/em\u003e and\u003cem\u003e P \u003c/em\u003evalues for gene quantifications based on the long reads (LRs) of ONT DRS and the short reads (SRs) of Illumina. Each dot represents a gene.\u003cstrong\u003e D: \u003c/strong\u003eHistogram showing the proportion of ONT LRs in the non-coding regions of the sequenced samples.\u003cstrong\u003e E:\u003c/strong\u003e Pie chart illustrating lincRNA annotations based on DRS LR that were supported by 5 SR reads.\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/925548cfbc9044c2f124afe3.jpg"},{"id":82206119,"identity":"d9092d3e-c5e6-4b02-a59c-a7274a9810c1","added_by":"auto","created_at":"2025-05-07 17:28:45","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":85902,"visible":true,"origin":"","legend":"\u003cp\u003eBoth long-read (LR) and short-read (SR) sequencing reveal substantial lincRNA activation in methylation-deficient mutants.\u003cstrong\u003e A:\u003c/strong\u003e Heatmap of lincRNA expression profiles obtained using long reads (LRs) for quantification.\u003cstrong\u003e B: \u003c/strong\u003eHeatmap of lincRNA expression profiles obtained using short reads (SRs) for quantification.\u003cstrong\u003e C:\u003c/strong\u003e Box plot comparing lincRNA expression levels in the different datasets. The left side is based on long reads and the right on short reads. \u003cstrong\u003eD: \u003c/strong\u003eHeatmap of PCG expression profiles obtained using LRs for quantification. Six hundred lincRNAs were randomly sampled.\u003cstrong\u003e E: \u003c/strong\u003eBox plot comparing protein-coding gene expression levels in the different datasets. The left side is based on long reads and the right on short reads. \u003cstrong\u003eF:\u003c/strong\u003e Venn diagram showing the coincidence of differentially expressed lincRNAs as determined using LR and SR data. \u003cstrong\u003eG:\u003c/strong\u003e IGV visualization of a \u003cem\u003eddm1-\u003c/em\u003eactivated lincRNA locus, \u003cem\u003eMSTRG.12166\u003c/em\u003e, using LR and SR libraries.\u003cstrong\u003eH:\u003c/strong\u003e IGV visualization of a lincRNA locus activated in both \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e, , using LR and SR libraries.\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/72a04b3c37c72ca8ad1093cf.jpg"},{"id":82206121,"identity":"ff2e6109-03a4-4107-9db5-ef24f7de2c65","added_by":"auto","created_at":"2025-05-07 17:28:45","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":83408,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in ddm1 mutants were widely detected and showed longer length. \u003cstrong\u003eA:\u003c/strong\u003e Histogram showing the ratio of transcript detected with m6A, m5C and poly(A) signals in different samples. \u003cstrong\u003eB: \u003c/strong\u003eDistribution of poly(A) length based on ONT sequencing among \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs, non-\u003cem\u003eddm1-\u003c/em\u003eactivated lincRNAs, and PCGs.\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/8ae5fdb52d6572bca3697dab.jpg"},{"id":82206379,"identity":"87c45bd9-3fbf-47a2-bb2c-bf70f42fa0c8","added_by":"auto","created_at":"2025-05-07 17:36:45","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":59876,"visible":true,"origin":"","legend":"\u003cp\u003eDecreased DNA methylation activates novel lincRNA transcription.\u003cstrong\u003e A: \u003c/strong\u003eBoxplot showing DNA methylation levels at CpG sites based on ONT DDS.\u003cstrong\u003e B:\u003c/strong\u003e Boxplot showing expression levels of\u003cstrong\u003e \u003c/strong\u003e\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs.\u003cstrong\u003e C:\u003c/strong\u003e Boxplot showing expression of non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs.\u003cstrong\u003e D:\u003c/strong\u003e Boxplot showing expression of PCGs in\u003cem\u003e ddm1\u003c/em\u003e mutants and wild-type as indicated at the bottom.\u003cstrong\u003e E: \u003c/strong\u003eIGV visualization of the \u003cem\u003eddm1-\u003c/em\u003eactivated lincRNA \u003cem\u003eMSTRG.12166\u003c/em\u003e. From top to bottom: ONT DRS data showing the expression of lincRNA \u003cem\u003eMSTRG.12166\u003c/em\u003e;\u003cstrong\u003e \u003c/strong\u003eDNA long-read sequencing (ONT) DDS showing the absence of structural variation in this lincRNA and the decrease of DNA methylation in \u003cem\u003eddm1\u003c/em\u003e.\u003c/p\u003e","description":"","filename":"Picture5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/55fb176852dc843e1de11a52.jpg"},{"id":82206383,"identity":"7f231da5-b898-4870-aff7-a587733f7a8b","added_by":"auto","created_at":"2025-05-07 17:36:46","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":67562,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs are widespread in natural populations at low frequency. \u003cstrong\u003eA: \u003c/strong\u003eBar plot showing the number of \u003cem\u003eddm1\u003c/em\u003e-activated and non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs expressed in natural populations.\u003cstrong\u003e B: \u003c/strong\u003eFrequency of expression (FPKM \u0026gt; 1) for \u003cem\u003eddm1\u003c/em\u003e-activated and non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in natural populations. The x-axis represents frequency. \u003cstrong\u003eC: \u003c/strong\u003eDistribution of the coefficient of variation (CV, ratio of standard deviation to mean) among \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs, non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs, and PCGs expressed in natural populations. \u003cstrong\u003eD:\u003c/strong\u003e IGV visualization of a \u003cem\u003eddm1\u003c/em\u003e-activated lincRNA, \u003cem\u003eMSTRG.19799\u003c/em\u003e, detected in eight accessions. \u003cstrong\u003eE: \u003c/strong\u003eDNA methylation levels of \u003cem\u003eMSTRG.19799 \u003c/em\u003ein selected \u003cem\u003eArabidopsis\u003c/em\u003e accessions.\u003c/p\u003e","description":"","filename":"Picture6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/28019ea60011565a442d08ff.jpg"},{"id":82206126,"identity":"fad793d2-1483-448c-9dd9-8e6493141f04","added_by":"auto","created_at":"2025-05-07 17:28:46","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":68522,"visible":true,"origin":"","legend":"\u003cp\u003eMachine learning predicts \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs.\u003cstrong\u003e A:\u003c/strong\u003e The three model perform the best are Random Forest, XGBoost and SVM Linear. The histogram shows mean accuracy and the dots represent the five-fold cross-validation results.\u003cstrong\u003e B:\u003c/strong\u003e Receiver operating characteristic (ROC) curve for prediction of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in current data.\u003cstrong\u003e C:\u003c/strong\u003e Precision-recall curve for prediction of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in current data.\u003cstrong\u003e D: \u003c/strong\u003eSchematic of the model. It takes as input the 100 bp DNA sequences containing the lincRNA and predicts transcriptional activity given deficient DNA methylation. DNA sequence contexts were converted to one-hot encoding and transformed into higher-level features in a convolutional neural network to make predictions. \u003cstrong\u003eE: \u003c/strong\u003eROC curve for prediction of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in current data. \u003cstrong\u003eF: \u003c/strong\u003ePrecision-recall curve for prediction of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in current data.\u003c/p\u003e","description":"","filename":"Picture7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/7fe6934e5898a77cb555c356.jpg"},{"id":102095042,"identity":"8301bcd4-4d3f-4f29-baae-9b4240c313f1","added_by":"auto","created_at":"2026-02-07 08:11:32","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1998948,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/467cede9-80af-4ea9-983e-6d3f30014d96.pdf"},{"id":82206380,"identity":"3bd2734f-c71d-4623-95ff-b18c3b15258e","added_by":"auto","created_at":"2025-05-07 17:36:46","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":269457,"visible":true,"origin":"","legend":"Supplementary Figure 1","description":"","filename":"FigS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/5623fe081aa2caff27c17d79.pdf"},{"id":82207188,"identity":"059a8a56-2218-4ae6-8eec-92921900d079","added_by":"auto","created_at":"2025-05-07 17:52:46","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":132417,"visible":true,"origin":"","legend":"Supplementary Figure 2","description":"","filename":"FigS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/e6307313d257803dbcc375d5.pdf"},{"id":82206390,"identity":"c5f9cf7d-dac6-433a-ace0-75f9b172761e","added_by":"auto","created_at":"2025-05-07 17:36:46","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":15336349,"visible":true,"origin":"","legend":"Supplementary Figure 3","description":"","filename":"FigS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/20712526f6bb0d6580edfff6.pdf"},{"id":82206132,"identity":"8c046e17-a30d-4733-9d3d-c6c7f91423d3","added_by":"auto","created_at":"2025-05-07 17:28:46","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":144665,"visible":true,"origin":"","legend":"Supplementary Figure 4","description":"","filename":"FigS4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/5d214c5a20d2c28347f3ea11.pdf"},{"id":82206135,"identity":"6afe855b-b2a2-41c5-a5b0-86061a7f12f3","added_by":"auto","created_at":"2025-05-07 17:28:46","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":2056927,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 5\u003c/p\u003e","description":"","filename":"FigS5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/2add9a7e46267233b03207c2.pdf"},{"id":82206861,"identity":"92ec9635-7e0f-447a-be0a-ae06555dbc15","added_by":"auto","created_at":"2025-05-07 17:44:46","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":114628,"visible":true,"origin":"","legend":"Supplementary Figure 6","description":"","filename":"FigS6.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/04f23263bcf2067567b30c88.pdf"},{"id":82206134,"identity":"cf900e42-0364-4e54-aa2d-93aaf15ca197","added_by":"auto","created_at":"2025-05-07 17:28:46","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":256461,"visible":true,"origin":"","legend":"Supplementary Figure 7","description":"","filename":"FigS7.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/02e3c355f6c269bb7ee53ce6.pdf"},{"id":82206385,"identity":"4147bd97-50fc-4445-9995-b393f1a1d1d8","added_by":"auto","created_at":"2025-05-07 17:36:46","extension":"xlsx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":229802,"visible":true,"origin":"","legend":"Supplementary Table","description":"","filename":"TableS.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/bb5834587725048f031d3891.xlsx"},{"id":82206386,"identity":"786181a5-6f1b-4474-b2d1-f6fd4b9a718a","added_by":"auto","created_at":"2025-05-07 17:36:46","extension":"docx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":732967,"visible":true,"origin":"","legend":"The description of supplementary files","description":"","filename":"DescriptionofSupplementaryFiles.docx","url":"https://assets-eu.researchsquare.com/files/rs-6399991/v1/592eac9b88e8633957294568.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Nanopore Direct RNA-Seq Reveals Widespread and Predictable Non-Coding Transcriptional Variations in DNA Methylation-deficient Arabidopsis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eNon-coding regions of genomes are reported to have great potential for transcription, and therefore play important roles in genomic evolution and regulation \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. Long intergenic non-coding RNAs (lincRNAs), which are transcribed from intergenic regions and encode no recognizable proteins, are widespread in eukaryotic transcriptomes and exert a wide spectrum of influences within an organism \u003csup\u003e\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Indeed, mounting evidence suggests that in plants, lincRNAs play regulatory roles in developmental and physiological processes including vernalization \u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, flowering development \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e, stress response \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e, and crop domestication \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. LincRNAs as an unclassified pool are far more diverse than protein-coding genes (PCGs) in terms of expression activity, increased tissue specificity, and variability in expression across individuals \u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Transcription of lincRNAs can result in novel genes and provide a route for expanding genome functionality without changing the number of genes \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. From an evolutionary perspective, the flexibility of lincRNAs is important for generating genomic and phenotypic plasticity, such as constructive neutral evolution.\u003c/p\u003e \u003cp\u003eIt has been widely proposed that epigenetic variations can affect gene expression and cause visible phenotypic differences, allowing rapid adaptation to dynamic environmental changes \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. Epigenome-wide association studies (EWAS) in natural populations along with studies using epigenetic recombinant inbred lines (epiRILs) have identified a large number of functional epigenetic variations \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. More importantly, populational genetic analysis and comparative genome analysis have shown that a proportion of lincRNAs are under positive selection pressure \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. That lincRNAs serve as a source of variation underlines the high potential of epigenome exploitation for breeding purposes.\u003c/p\u003e \u003cp\u003eA relationship between lincRNAs and DNA methylation has been widely reported in animals and plants at both individual and population levels\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. For example, population studies in the model plant \u003cem\u003eArabidopsis\u003c/em\u003e and in cancer cells have reported that lincRNA activation is associated with hypomethylation \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e. In plants, reprogramming of DNA methylation in the course of hybridization and polyploidization are reported to lead to accumulation of non-coding RNAs \u003csup\u003e\u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Studies of DNA methylation-related mutants have also characterized a large number of lincRNAs to be activated in multiple plant and animal species, including \u003cem\u003eOryza sativa\u003c/em\u003e, \u003cem\u003eArabidopsis thaliana\u003c/em\u003e, \u003cem\u003eZea mays\u003c/em\u003e, \u003cem\u003eBrassica rapa\u003c/em\u003e, and \u003cem\u003eMus musculus\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e,\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Taken together, these studies indicate that genomes have a capability to integrate epigenetic modifications and transcription machinery on non-coding regions to create novel transcripts as a candidate pool for adaptive evolution within a short time. It is thus of great interest to identify and validate at genomic scale the flexibility of those non-coding regions having transcription potential.\u003c/p\u003e \u003cp\u003eIn \u003cem\u003eArabidopsis\u003c/em\u003e, DNA methylation mutants exhibiting different levels of DNA methylation deficiency have been well characterized \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. Nucleosome remodelers of the \u003cem\u003eDNA methylation 1\u003c/em\u003e (\u003cem\u003eDDM1\u003c/em\u003e)/Lsh family, which encode a protein similar to the chromatin-remodeling factor SWI2/SNF2, are required for maintenance of CG methylation (mCG) that is conserved in plants and animals \u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Meanwhile, the protein encoded by \u003cem\u003emethyltransferase 1\u003c/em\u003e (\u003cem\u003eMET1\u003c/em\u003e) recognizes hemimethylated CG following DNA replication and methylates the naked cytosine in the daughter strand \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. Subsequent maintenance of DNA methylation in heterochromatin requires \u003cem\u003eDDM1\u003c/em\u003e \u003csup\u003e30\u003c/sup\u003e. Accordingly, \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e mutations have been shown to result in drastically decreased DNA methylation spreading on transposon elements (TEs) \u003csup\u003e\u003cspan additionalcitationids=\"CR32\" citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. In addition, the hypomethylated status characteristic of \u003cem\u003eddm1\u003c/em\u003e can be inherited even after eight generations \u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. These features make \u003cem\u003eArabidopsis ddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e mutants excellent models for the investigation of lincRNA behaviors under a precondition of DNA hypomethylation \u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020)(Tao et al., 2020).\u003c/p\u003e \u003cp\u003eAbout 80% of lincRNAs have 5\u0026prime; caps and 3\u0026prime; poly(A) tails, making them amenable for discovery by oligo (dT)-based RNA sequencing \u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e. However, lincRNA sequences are also enriched with transposable element (TE) fragments \u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e,\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e, which highly repetitive sequences cannot be accurately assembled by the next-generation sequencing (NGS, i.e., Illumina) technologies on which current studies rely; such technologies lack the necessary long-range information \u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e,\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e. Direct RNA sequencing (DRS) offered by Oxford Nanopore Technologies (ONT) can sequence polyadenylated RNAs directly, without the recoding and amplification biases inherent to other sequencing methodologies \u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. Moreover, DRS can produce long reads (up to 30 kb, vs up to 600 bp at present for Illumina RNA-seq), which can provide for accurate transcript assemblies \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e. In addition, this method can detect the RNA modifications of m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA and m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC simultaneously, which introduce further epi-modification features \u003cem\u003ein situ\u003c/em\u003e at single-molecule resolution \u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e,\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. Therefore, ONT DRS is an advanced technology suitable for the detailed study of lincRNAs.\u003c/p\u003e \u003cp\u003eTo investigate the underrepresented complexity of non-coding elements in the transcriptome, we leveraged the power of ONT DRS to characterize lincRNAs in \u003cem\u003eArabidopsis\u003c/em\u003e lines having null mutations of the DNA methylation regulators \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e. This approach was supplemented with short reads produced by Illumina RNA-seq.\u0026nbsp;A large number of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs were identified and their occurrence in a natural population dissected. In addition, random forest classification algorithms were applied to directly learn DNA sequences traits from lincRNAs for the prediction of lincRNA sensitivity to activation. The presented models can serve as practical screening tools for mining causative non-coding variations in natural and artificial populations.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy design\u003c/h2\u003e \u003cp\u003eLeaves from \u003cem\u003eArabidopsis thaliana\u003c/em\u003e Columbia (Col-0) wild type (WT) and strains respectively having mutations in \u003cem\u003eDDM1\u003c/em\u003e (TAIR Germplasm name: \u003cem\u003eddm1-2\u003c/em\u003e, AT5G66750) and \u003cem\u003eMET1\u003c/em\u003e (TAIR Germplasm name: CS16394, \u003cem\u003emet1-3\u003c/em\u003e AT5G49160) (\u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e) were prepared and sequenced using RNA Direct Sequencing, ONT DNA sequencing, and Illumina Strand-Specific RNA sequencing (ssRNA-seq).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePlant materials and growth conditions\u003c/h3\u003e\n\u003cp\u003e \u003cem\u003eArabidopsis thaliana\u003c/em\u003e Col-0 WT, \u003cem\u003eddm1\u003c/em\u003e, and \u003cem\u003emet1\u003c/em\u003e seeds were surface sterilized with 20% sodium hypochlorite (NaClO) for 10 min, followed by 70% ethanol for 5 min. Seeds were then washed four times in sterile distilled water, plated on growth medium, and placed in a growth chamber with 16 h light/8 h dark for 30 days. Thirty-day leaves (2 g) were harvested with four biological replicates and immediately immersed in liquid nitrogen. Two samples were used for an ONT DRS pilot run with a single flow cell (R9.4.1), while the remaining samples were used for ssRNA-seq.\u003c/p\u003e\n\u003ch3\u003eNanopore direct RNA sequencing library preparation\u003c/h3\u003e\n\u003cp\u003eTotal RNA from \u003cem\u003eArabidopsis\u003c/em\u003e leaves was isolated with Trizol according to the manufacturer\u0026rsquo;s protocol, and then polyA-selected using oligodT Dynabeads (Invitrogen). Approximately 200\u0026ndash;500 ng of poly(A)-tailed mRNA was purified using a Dynabeads mRNA Purification Kit, followed by library construction using a Direct RNA Sequencing Kit. The libraries were sequenced on MinION MkIb instruments with a R9.4 sequencing microarray over 48\u0026ndash;72 h using PromethION (Oxford Nanopore Technologies Ltd.). Each biological replicate and condition were sequenced independently in a different flow cell. Sequencing was performed by Wuhan Benagen Tech Solutions Company Limited, and the recommended scripts in MinKNOW were used to generate multi-fast5 read files with live base calling enabled.\u003c/p\u003e\n\u003ch3\u003eAlignment of ONT DRS\u003c/h3\u003e\n\u003cp\u003eMulti-fast5 reads were base-called using Guppy (v 0.1.10) with the default parameters \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e, which can achieve a calling accuracy of more than 99% \u003csup\u003e45\u003c/sup\u003e. LRs were separated into pass and fail folders; only reads in the pass folder (Phred\u0026thinsp;\u0026gt;\u0026thinsp;7) were used in subsequent analysis. For each sample, ONT LRs were aligned to the \u003cem\u003eArabidopsis thaliana\u003c/em\u003e TAIR10 reference genome using Minimap2 (v 2.18, github.com/lh3/minimap2)), a long read local splice-aware aligner, with the following parameters: -ax splice -k14 -uf --secondary\u0026thinsp;=\u0026thinsp;no \u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e. The resulting alignment files were converted to BAM format, sorted, and indexed by sequence coordinate with SAMtools (v1.6) \u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e. Aligned reads with a secondary or supplementary aligned flag were discarded. Only reads with alignment flags 0 and 16 (primary mapping, forward strand and reverse strand of genome) were retained for further analysis. ONT reads were also aligned to the transcriptome (combined cDNA and non-coding RNA reference FASTA files) using Minimap2 (v 2.1)with the following parameters: -ax map-ont \u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003ch3\u003eGeneration of lincRNA annotations using ONT DRS\u003c/h3\u003e\n\u003cp\u003eFor RNA annotation, we combined primary genome alignments from all libraries in each ONT data set, then assembled transcripts using StringTie2 (v2.0) in long read mode with the parameters -L \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e,\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e.The identified transcripts from each data set were compared with the set of annotated transcripts using StringTie2 (V2.0) \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. For better annotation of isoforms, reads with similar splice junctions were grouped to generate a raw isoform annotation using FLAMES \u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. FLAMES can correct potential splice site and transcript start/end errors by comparing the raw isoform annotation against the reference annotation. Transcripts that have splice junctions (\u0026lt;\u0026thinsp;5 bp) and transcript start/end (\u0026lt;\u0026thinsp;100 bp) similar to the reference transcript were merged with the reference; thus, incomplete transcripts were collapsed. High-confidence isoforms were those identified with at least five supporting reads. Finally, we created a GTF file of the transcripts (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ZhaoTingzju/dRNA-seq\u003c/span\u003e\u003cspan address=\"https://github.com/ZhaoTingzju/dRNA-seq\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Novel intergenic transcripts annotated with class code \u0026ldquo;u\u0026rdquo; by cuffcompare \u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e were considered potential lincRNAs. The sequence of each transcript was extracted from its assembled form, and coding potential was assessed with CPC2 (v2.0b) \u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e and Pfam_scan using HMMER (v 3.2.1, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.hmmer.org/\u003c/span\u003e\u003cspan address=\"http://www.hmmer.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The obtained LR-RNA annotation, including coding genes, non-coding RNAs, and both novel and known annotated transcripts, are available online (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ZhaoTingzju/dRNA-seq\u003c/span\u003e\u003cspan address=\"https://github.com/ZhaoTingzju/dRNA-seq\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eGene and transcript abundance estimation\u003c/h2\u003e \u003cp\u003eLR were aligned to the LR-RNA annotation and expression at gene level was quantified using StringTie2 (v2.0) \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. To quantify genes, mapped reads were assigned to individual genes and counted by FeatureCount (v1.6.2) \u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e with the arguments LongRead\u0026thinsp;=\u0026thinsp;TRUE (as recommended in the program help page when dealing with Nanopore data) and primaryOnly\u0026thinsp;=\u0026thinsp;TRUE to count primary alignments only. Differential gene expression analysis was conducted using the DESeq2 package\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e with the negative binomial statistic approach. The cutoff criteria for the selection of differentially expressed genes (DEGs) were: | log2-fold change | \u0026gt; 1 and false discovery rate (FDR)\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePoly(A) tail length estimation and alternative polyadenylation using ONT DRS\u003c/h3\u003e\n\u003cp\u003eONT DRS sequencing requires a poly(A) tail for first adapter ligation. Poly(A) tail length was estimated from raw signal for each read using the poly(A) estimation function of the program Nanopolish (v0.11.0) \u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e with \u0026ldquo;polya-merged\u0026rdquo; branch. Length estimates were only considered if the QC tag was Pass \u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e, and were subsequently grouped by gene and isoform using an in-house Perl script.\u003c/p\u003e\n\u003ch3\u003eDetection of methylated adenosine and cytosine in ONT DRS\u003c/h3\u003e\n\u003cp\u003eBase-called multi-fast5 reads were converted to single-read fast5s using the Oxford Nanopore Technologies API ont_fasta5 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/nanoporetech/ont_fast5_api\u003c/span\u003e\u003cspan address=\"https://github.com/nanoporetech/ont_fast5_api\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). RNA modifications were identified using the software Tombo (v1.5.1) \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. First, raw signal squiggles were assigned to the reference sequence using \u003cem\u003eresquiggle\u003c/em\u003e. Next, modified base detection was carried out using \u003cem\u003edetect modifications.\u003c/em\u003e The models of \u0026ldquo;5mC\u0026rdquo; and \u0026ldquo;\u003cem\u003ede novo\u003c/em\u003e\u0026rdquo; were implemented separately to detect possible modifications in each read. MINES was used to detect m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA with default parameters (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/YeoLab/MINES\u003c/span\u003e\u003cspan address=\"https://github.com/YeoLab/MINES\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eONT DDS library preparation\u003c/h2\u003e \u003cp\u003eHigh-quality genomic DNA was isolated from fresh leaves using the CTAB method, and DNA quality and concentration were confirmed using 0.75% agarose gel electrophoresis, a NanoDrop One spectrophotometer (Thermo Fisher Scientific), and a Qubit 3.0 Fluorometer (Life Technologies, Carlsbad, CA, USA). Libraries were prepared with the SQK-LSK109 ligation kit according to the standard protocol. Each purified library was loaded onto a primed R9.4 Spot-On Flow Cell and sequenced using a PromethION sequencer (Oxford Nanopore Technologies, Oxford, UK) with 48-h runs at Wuhan Benagen Tech Solutions Company Limited, Wuhan, China. Base calling analysis of raw data was performed using the Oxford Nanopore Guppy software (v 0.1.10) \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eStructure variation calling based on ONT DDS\u003c/h2\u003e \u003cp\u003eMulti-fast5 reads were base-called using Guppy (v 0.1.10) with the default parameters \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Passing LR reads were aligned to the \u003cem\u003eArabidopsis thaliana\u003c/em\u003e TAIR10 reference genome with Minimap2 (v 2.1) \u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e, and SVs were called with cuteSV (v1.0.12), a sensitive SV detection implementation for LRs, from processed bam files using the suggested parameters for ONT: --max_cluster_bias_INS 100, --diff_ratio_merging_INS 0.3, -max_cluster_bias_DEL 100, --diff_ratio_merging_DEL 0.3 \u003csup\u003e57\u003c/sup\u003e. SV information was stored in VCF format \u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ZhaoTingzju/dRNA-seq\u003c/span\u003e\u003cspan address=\"https://github.com/ZhaoTingzju/dRNA-seq\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eIllumina short-read RNA sequencing library preparation\u003c/h2\u003e \u003cp\u003eStrand-specific RNA-seq libraries were constructed following the TruSeq RNA sample preparation protocol as previously described with modifications \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. Total RNA was treated with DNase first and then depleted of rRNA using complementary DNA oligos (IDT) and RNase H (Invitrogen, Waltham, MA, USA). Afterwards, RNA sequencing was conducted on the Illumina Hiseq 2500 platform with 150-bp paired-end reads according to the manufacturer\u0026rsquo;s protocol. Sequencing was performed by Wuhan Benagen Tech Solutions Company Limited.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eGene and transcript abundance estimation based on SRs\u003c/h2\u003e \u003cp\u003eFor SR-RNA annotation, the strand-specific paired-end 150-bp reads were first quality trimmed using Fastp \u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e, then aligned to the \u003cem\u003eArabidopsis thaliana\u003c/em\u003e TAIR10 genome using Hisat2 with the following parameter: --dta \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. Transcripts were subsequently assembled using StringTie (v2.0) with the \u0026ndash;dta flag \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. Differential gene expression analysis was conducted using the DESeq2 package with the negative binomial statistic approach \u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e. Cutoff criteria for the selection of differentially expressed genes (DEGs) were: | log2-fold change | \u0026gt; 1 and false discovery rate (FDR)\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of publicly available data\u003c/h2\u003e \u003cp\u003ePublicly available SR transcriptomes from 675 diverse \u003cem\u003eArabidopsis thaliana\u003c/em\u003e accessions from the 1001 Genomes Project were downloaded from Gene Expression Omnibus (GEO, accession id: GSE80744). The reads were quality trimmed using Fastp \u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e, then aligned to the \u003cem\u003eArabidopsis thaliana\u003c/em\u003e TAIR10 genome using Hisat2 with the following parameter: --dta \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. Transcripts were subsequently quantified using StringTie (v2.0) based on the LR-RNA annotation \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. The resulting data matrix, which contained expression values (FPKM) for all samples, was deposited in GitHub (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ZhaoTingzju/dRNA-seq\u003c/span\u003e\u003cspan address=\"https://github.com/ZhaoTingzju/dRNA-seq\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eCorresponding methylomes were likewise retrieved from GEO (accession id: GSE54292). The Illumina BS-Seq data was quality-checked with Fastp \u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e. The clean reads were then aligned to the corresponding reference, and DNA methylation levels were calculated using Bismark \u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003cb\u003ePrediction of\u003c/b\u003e \u003cb\u003eddm1\u003c/b\u003e\u003cb\u003e-activated lincRNAs\u003c/b\u003e\u003c/p\u003e \u003cp\u003eWe implemented six machine learning models including Random Forest, XGBoost, SVM Linear, SVM RBF, Nearest Neighbors and Naive Bayes Classifier by using the library scikit learn in Python 3.10.9. The optimal parameter combination was selected based on algorithm accuracy as determined by five-fold cross-validation (CV). The scripts are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ZhaoTingzju/dRNA-seq\u003c/span\u003e\u003cspan address=\"https://github.com/ZhaoTingzju/dRNA-seq\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eIn this study, we generated three broad categories of features: sequence features, gene structure and expression features, and RNA modification features. The sequence properties included the ratios of single nucleotides and dinucleotides, sequence length. Also, we upload sequence data on CPC2 (v2.0b) \u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e to get more information including predicted coding probability and other features. TE annotation file was downloaded from TAIR10 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.arabidopsis.org\u003c/span\u003e\u003cspan address=\"http://www.arabidopsis.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to obtain the feature of the TE number in each lncRNA.\u003c/p\u003e \u003cp\u003eThe convolutional neural network was also implemented by using the library scikit learn in Python 3.10.9. The optimal parameter combination was selected based on algorithm accuracy as determined by ten-fold cross-validation. DNA sequences of both active and non-active lincRNAs were used as training input. the numbers of positive (\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs) and negative (non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs) labels in the dataset are set to be same.\u003c/p\u003e \u003cp\u003eSequences of length 100 bp was converted into a 2-D matrix of size 4 \u0026times; L, with each column a one-hot vector encoding the presence of the four DNA nucleotides A, C, G, and T. Method performance was evaluated using the ROC. The scripts are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ZhaoTingzju/dRNA-seq\u003c/span\u003e\u003cspan address=\"https://github.com/ZhaoTingzju/dRNA-seq\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eStatistics\u003c/h2\u003e \u003cp\u003eStatistical analysis was performed using Python 3.10.9 and R 4.4.1 software.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003e \u003cb\u003eWorkflow of DRS and lincRNA annotation for\u003c/b\u003e \u003cb\u003eArabidopsis thaliana\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo investigate the fine details of lincRNAs, including expression variations, RNA modifications, and alternative poly(A) tails (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA), ONT DRS and Direct DNA Sequencing were performed on DNA isolated from leaf samples of \u003cem\u003eArabidopsis thaliana\u003c/em\u003e Col-0 wild type (WT) and \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e mutants. Illumina Strand-Specific RNA sequencing (ssRNA-seq) with 150-bp short reads (SRs) carried out in parallel on the same samples was also conducted for comparsion (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe analysis workflow for ONT DRS is shown in Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. Base-calling and adapter trimming of the ONT long reads (LRs) was performed using Guppy (v0.1.10) \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e and yielded 41\u0026nbsp;million passing reads (Phred quality score\u0026thinsp;\u0026gt;\u0026thinsp;7) from the six samples (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). Taking advantage of the unprecedented read length of ONT DRS LRs, we assembled them for \u003cem\u003ede novo\u003c/em\u003e RNA annotation. The analysis workflow is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB. LRs from all libraries were combined for transcript assembling using the LR model of StringTie2 (v2.0) \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e,\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. Transcripts with similar splice junctions (\u0026lt;\u0026thinsp;5 bp by default) and transcript start/end positions (\u0026lt;\u0026thinsp;100 bp) were further merged using FLAMES (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/LuyiTian/FLAMES\u003c/span\u003e\u003cspan address=\"https://github.com/LuyiTian/FLAMES\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and those having at least five mapped LRs were retained.\u003c/p\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eAssessment of DRS profiles and lincRNA annotation result\u003c/h2\u003e \u003cp\u003eThe analysis of ONT DRS produced abundant high-quality reads. The median length of passed reads is 893 bp, with the maximum length being 14,248 bp. The read length distribution and read quality of LRs were similar to values reported in a previously published study of \u003cem\u003eArabidopsis thaliana\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e, but the total reads obtained was above 6\u0026nbsp;million (M) on average, which improved the total base coverage up to 7.4-fold more than the previous study, at about 0.8 M per library \u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe LRs were mapped to the \u003cem\u003eArabidopsis thaliana\u003c/em\u003e reference genome TAIR10 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.arabidopsis.org\u003c/span\u003e\u003cspan address=\"http://www.arabidopsis.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and associated RNA annotations by Minimap2 (v2.1) \u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e, yielding average mapping rates of 99.70% and 96.85%, respectively (Fig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e; Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). The sequencing depth of each data set were further assessed with the frequencies of gene number identified and expression quantification using LRs. Although the total number of genes detected when using LRs was relatively less than the tally from SRs, at least 12,000 genes were efficiently evaluated by LRs (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). LRs were randomly subsampled in amounts ranging from 0.5 to 5 M reads, and detections of expressed genes reached plateau phase at about 2.5 M reads (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB). Notably, read counts obtained from LR and SR sequencing were highly correlated, with Spearman\u0026rsquo;s correlation coefficient values ranging from 0.84 to 0.88 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC; Fig. \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e). Taken together, the profiles of the obtained ONT DRS data were consistent and allowed quantitative analysis of transcriptome.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWith regard to non-coding transcripts, less than 1% of LRs in the WT mapped to non-coding regions, but the amounts of non-coding reads were six-fold higher in \u003cem\u003eddm1\u003c/em\u003e and two-fold higher in \u003cem\u003emet1\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). This trend is in agreement with a previously reported observation of non-coding RNA activation in DNA methylation mutants \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e, which suggests that the LR libraries contained reliable RNAs from non-coding regions.\u003c/p\u003e \u003cp\u003eThe obtained LR-RNA annotation comprised 42,305 transcripts representing 15,925 genes. LincRNAs were further annotated by calculating coding potential and evaluating the presence of domains using coding potential calculator2 (CPC2, v2.0b) \u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e and HMMER (v 3.2.1, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://hmmer.org/\u003c/span\u003e\u003cspan address=\"http://hmmer.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Transcripts found to have protein coding potential or TE domains were excluded. Ultimately, 1,663 transcripts from 549 transcription models were annotated as lincRNAs. Of those, 98% were also represented in SR libraries (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eE), confirming that ONT DRS can effectively capture the complex body of native full-length lincRNAs.\u003c/p\u003e \u003cp\u003e \u003cb\u003eActive expression of lincRNAs in\u003c/b\u003e \u003cb\u003eddm1\u003c/b\u003e \u003cb\u003eand\u003c/b\u003e \u003cb\u003emet1\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo examine the impact of DNA methylation on lincRNA transcriptional activity, PCGs and lincRNAs were quantified using normalized fragments per kilobase per million mapped reads (FPKM). For both LR and SR libraries, lincRNA expression levels in \u003cem\u003eddm1\u003c/em\u003e (LR median\u0026thinsp;=\u0026thinsp;6.23; SR median\u0026thinsp;=\u0026thinsp;2.64) and \u003cem\u003emet1\u003c/em\u003e (LR median\u0026thinsp;=\u0026thinsp;0.87; SR median\u0026thinsp;=\u0026thinsp;0.43) were obviously higher than those in the WT (LR median\u0026thinsp;=\u0026thinsp;0; SR median\u0026thinsp;=\u0026thinsp;0) (typically \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026le;\u0026thinsp;2 \u0026times; 10\u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e, Wilcoxon test) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA-\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). In contrast, the deficiency of DNA methylation in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e mutants affected only a small fraction of PCG transcription in both LR and SR libraries, although the variations also reached statistical significance (typically \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026le;\u0026thinsp;2 \u0026times; 10\u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e, Wilcoxon test) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD and \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). Trends in active lincRNAs in the hypomethylated mutants were confirmed by calculating the number of differentially expressed genes (DEGs). Of the annotated lincRNAs, 61.57% (n\u0026thinsp;=\u0026thinsp;338) and 22.22% (n\u0026thinsp;=\u0026thinsp;122) were differentially up-regulated in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e, respectively (adjusted \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). For PCGs, the corresponding ratios were 2.41% (n\u0026thinsp;=\u0026thinsp;662) and 1.55% (n\u0026thinsp;=\u0026thinsp;426) in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e, respectively. The same DEG analysis workflow was also applied to the SR libraries, and the majority of lincRNA DEGs were found to be shared between LR and SR libraries. Among the up-regulated lincRNAs detected in LR libraries, 310 (91.72%, out of 338) and 116 (95.08%, out of 122) were also differentially expressed in SR libraries in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eF and S3).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA representative lincRNA, \u003cem\u003eMSTRG.12166\u003c/em\u003e, was visualized in both LR and SR libraries using the Integrative Genomics Viewer (IGV) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eG). This transcript was specifically expressed in \u003cem\u003eddm1\u003c/em\u003e, but was not expressed in the corresponding WT or \u003cem\u003emet1\u003c/em\u003e samples. Similarly, Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eH illustrates a second representative lincRNA, \u003cem\u003eMSTRG.12270\u003c/em\u003e, that was actively expressed in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e but not expressed in WT.\u003c/p\u003e \u003cp\u003eAll told, the ONT DRS profiling confirmed our previous conclusions regarding the impact of DNA hypomethylation on activation of lincRNAs in multiple organisms, which might alter transcriptome plasticity \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. Thus, the lincRNAs were classified according to expression specificity, yielding 340 \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs and 209 non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs for further analysis (Table \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003ePatterns of RNA modifications and poly(A) tails on lincRNAs\u003c/h2\u003e \u003cp\u003eThe impacts of DNA hypomethylation on lincRNAs were further examined in relation to RNA m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA, m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC modification and poly(A) tail length. In total, an average of 37,808 high-confidence m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA loci were identified using Tombo (v 1.5.1) \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e and MINES \u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e. Each transcript had on average four m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA sites, which is consistent with a previous result in \u003cem\u003ePopulus trichocarpa\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e. Meanwhile, an average of 1,954,236 m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC loci were detected using Tombo (v 1.5.1) \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. In the \u003cem\u003eddm1\u003c/em\u003e mutant, the ratios of transcripts with m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA modification to those with m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC was similar between PCGs (m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA 68.90%; m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC 67.64%) and \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs (m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA 67.65%; m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC 63.83%), and slightly lower in non-\u003cem\u003eddm1\u003c/em\u003e activated lincRNAs (m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA 36.85%; m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC 42,59%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePoly(A) tails are important for transcription and transcript stability \u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e,\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e. Here we examined poly(A) tail length using the low-variance ionic current signal associated with the 3\u0026rsquo; end in ONT DRS \u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e. This method identified 15,014 transcripts with poly(A) tails (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA), which accounted for 95.10% of PCGs, 98.53% of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs, and 93.78% of non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs. Overall, non-coding RNAs featured longer 3\u0026rsquo; UTRs (\u003cem\u003eddm1\u003c/em\u003e-activated, mean 79 nt; non-\u003cem\u003eddm1\u003c/em\u003e-activated, mean 82 nt) than did PCGs (mean 68 nt) (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026le;\u0026thinsp;2 \u0026times; 10\u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e, Wilcoxon test) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB). The length of poly(A) tails was negatively correlated with transcriptional activity (Fig. \u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e). consistent with a previous report that poly(A) length negatively correlates with expression \u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eDNA methylation according to ONT DDS is negatively correlated with transcriptional activity at lincRNA loci\u003c/h2\u003e \u003cp\u003eTo investigate the quantitative relationships between DNA methylation and lincRNA transcriptional activity, the ONT DDS technology was employed to quantify DNA methylation levels without chemical treatment and structure variation. LR DNA sequencing with the ONT platform was performed using the same plant materials. Base calling was carried out using Guppy (v 0.1.10) [32], and yielded 3.10 M reads (Phred score\u0026thinsp;\u0026gt;\u0026thinsp;7) with an average length of 14 kb (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). The longest read obtained was 208,904 bp. The mean coverage of each sample was 61\u0026times; across the genome. The reads were mapped to the \u003cem\u003eArabidopsis thaliana\u003c/em\u003e genome TAIR10 by Minimap2 (v 2.1) [34], which revealed 189 structural variations in \u003cem\u003eddm1\u003c/em\u003e and 118 in \u003cem\u003emet1\u003c/em\u003e mutants in compare with WT (Fig. \u003cspan refid=\"MOESM6\" class=\"InternalRef\"\u003eS6\u003c/span\u003e). Structural variations were observed in only small fractions of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs (3.82%, n\u0026thinsp;=\u0026thinsp;13) and of non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs (2.87%, n\u0026thinsp;=\u0026thinsp;6); accordingly, there was no statistical enrichment of structural variation in \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.36, Fisher's Exact Test). (Table \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo examine the association of lincRNA activity with DNA methylation level, genome-wide DNA methylation was first confirmed in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;2.22 \u0026times; 10 \u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e) in compared with WT (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA) based on the calling of mC (methylated cytosine) frequency in reads by Nanopolish \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Comparison of gene body DNA methylation in \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs between \u003cem\u003eddm1\u003c/em\u003e (mean: 11.98%) and WT (mean: 79.17%) revealed a 67.20% decrease in the mutant (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB). Meanwhile, non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs and PCGs exhibited respective decreases of only 16.97% (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eC) and 5.58% (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD). These trends confirmed that the activation of lincRNAs in \u003cem\u003eddm1\u003c/em\u003e was significantly correlated with alteration of DNA modifications. This is further illustrated by a representative \u003cem\u003eddm1\u003c/em\u003e-activated lincRNA, \u003cem\u003eMSTRG.12166\u003c/em\u003e, at which locus the genomic sequences in WT and \u003cem\u003eddm1\u003c/em\u003e were identical, but RNA activation was specific to the \u003cem\u003eddm1\u003c/em\u003e mutant and corresponded with reduced DNA methylation (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eE). Hence, CG methylation is most likely a major obstacle to the expression of lincRNAs.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eddm1\u003c/b\u003e \u003cb\u003e-activated lincRNAs can be detected in\u003c/b\u003e \u003cb\u003eArabidopsis\u003c/b\u003e \u003cb\u003enatural populations\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe \u003cem\u003eArabidopsis ddm1\u003c/em\u003e mutant expresses a group of activated lincRNAs that have structural and epi-modification features similar to PCGs (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA); this finding indicates that non-coding regions of the genome might have the potential to generate new genes when released from DNA methylation. This hypothesis inspired us to examine the expression of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in natural populations so as to confirm the deduced lincRNA behavioral patterns at large scale.\u003c/p\u003e \u003cp\u003eTo further broaden the biological scope of this analysis, we obtained from a public database the leaf transcriptome profiles of 728 \u003cem\u003eArabidopsis\u003c/em\u003e natural accessions from the 1001 Genomes Project, along with accompanying original information concerning geographic distribution \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e (Table \u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e). We then investigated the expression of \u003cem\u003eddm1\u003c/em\u003e-activated and non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in those accessions, with expression defined as having FPKM\u0026thinsp;\u0026gt;\u0026thinsp;1. Interestingly, the majority of annotated lincRNAs (68.31%, n\u0026thinsp;=\u0026thinsp;375) were expressed in at least one of the tested accessions (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA). In addition, \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs originally identified under condition of hypomethylation were found to be active in this natural population, although only at low frequency (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eB). The expression variation of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs was significantly larger than that of either non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs or PCGs, determined by calculating the average coefficient variation (CV, \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026le;\u0026thinsp;2 \u0026times; 10\u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e, Wilcoxon test) (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eC). For example, one \u003cem\u003eddm1\u003c/em\u003e-activated lincRNA with multiple exons, \u003cem\u003eMSTRG.19779\u003c/em\u003e, was not expressed in most accessions (FPKM\u0026thinsp;\u0026lt;\u0026thinsp;1), similar to its state in Col-0 WT. However, it was highly expressed in at least three accessions, IP-San-10, Gd-1, Corig-1, similar to the situation in \u003cem\u003eddm1\u003c/em\u003e, with a high level of expression and a stable RNA splicing pattern (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eD). Moreover, the region showed significant decrease of DNA methylation in those accessions, including CG, CHG, CHH methylations (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eE). These results suggest that in natural populations, many accessions also present activation of lincRNA expression similar to what occurred in the \u003cem\u003eddm1\u003c/em\u003e mutant with genome-wide DNA hypomethylation.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003ePrediction of\u003c/b\u003e \u003cb\u003eddm1\u003c/b\u003e\u003cb\u003e-activated lincRNAs by machine learning and deep learning\u003c/b\u003e\u003c/p\u003e \u003cp\u003eAs noted above, \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs differ from non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs in terms of expression levels (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC) as well as epigenetic landscapes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). These patterns provide a fundamental basis for classifying lincRNAs into distinct groups. We then hypothesized that the observed activation patterns of lincRNAs are predictable, and through sufficient information, they can be classified into \u003cem\u003eddm1\u003c/em\u003e-activated or non-\u003cem\u003eddm1\u003c/em\u003e-activated categories.\u003c/p\u003e \u003cp\u003eThus, thirty-five features per lincRNA were constructed and used in six machine learning classification models including Random Forest, XGBoost (eXtreme Gradient Boosting), SVM Linear (Support Vector Machine with Linear Kernel), SVM RBF (Support Vector Machine with Radial Basis Function Kernel), Nearest Neighbors and Naive Bayes Classifier (Table \u003cspan refid=\"MOESM6\" class=\"InternalRef\"\u003eS6\u003c/span\u003e). Model performance was evaluated through five iterations of five-fold cross-validation and the result showed that Random Forest, XGBoost and SVM Linear models reached the best three prediction accuracy (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA and S7). Among these models, Random Forest achieved the highest AUC-ROC of 0.83 and average precision of 0.88, followed by XGBoost (AUC-ROC\u0026thinsp;=\u0026thinsp;0.81, AP\u0026thinsp;=\u0026thinsp;0.86) and SVM Linear (AUC-ROC\u0026thinsp;=\u0026thinsp;0.71, AP\u0026thinsp;=\u0026thinsp;0.76) (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB and \u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC). In addition, we ranked all features within each model in order of importance and presented the top seven features (Fig. \u003cspan refid=\"MOESM7\" class=\"InternalRef\"\u003eS7\u003c/span\u003eD). Notably, the number of transposable elements (TE) consistently ranked among the top two in the best three models (Random Forest, XGBoost and SVM Linear), indicating a potential relationship between TE abundance and lincRNA activation (Fig. \u003cspan refid=\"MOESM7\" class=\"InternalRef\"\u003eS7\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eInspired by the successful application of deep learning in direct DNA sequence, a deep convolutional neural network (DCNN) was then implemented for classifying lincRNAs as \u003cem\u003eddm1\u003c/em\u003e-activated or not based on the 100 bp sequence context (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eD). The performance of this fully sequence-based model was evaluated using a 10 times five-fold cross-validation procedure, in which it achieved an area under receiver operating characteristic value of 0.91, and average precision 0.91 (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eE and \u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eF). The DCNN model demonstrated better accuracy in the classification of \u003cem\u003eddm1\u003c/em\u003e-activated RNA compared to conventional machine learning models. This capability stems from DCNN's hierarchical architecture, where convolutional layers directly and automatically extracted multi-scale features related to lncRNA transcriptional activation from DNA sequences.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn this study, we demonstrated that ONT DRS technology is efficient for identifying at high confidence lincRNA transcript structures that have been overlooked by short-read sequencing. LincRNAs were finely mapped and epi-modifications examined at single molecule resolution. The \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs can also be activated in natural populations. Using thirty-five features, six machine learning models were constructed among which Random Forest model enabled prediction of transcript activation in the absence of DNA methylation with 70% accuracy. All told, with the integration of lincRNA activation patterns in populations and their predictable sequence patterns identified with machine-learning technology, we propose the use of lincRNAs and associated DNA methylation patterns as new molecular markers in the development of crop breeding strategies.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003eAdvantages of long-read sequencing technology in investigating non-coding RNAs\u003c/h2\u003e \u003cp\u003eAt present, ONT DRS is mainly applied in studies of protein coding genes to identify structure variations and epigenetic modifications based on its long-read features and the accessorized detection of RNA modifications. We demonstrate that ONT DRS can also be utilized to identify and quantify lincRNAs with high accuracy. It is known that lincRNAs are transcribed from TE-enriched intergenic regions \u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e, which feature repetitive sequences and genomic SVs associated with TE mobilizations; this makes it challenging to use short-read technology to navigate the structure of lincRNAs on both RNA and DNA levels.\u003c/p\u003e \u003cp\u003eThe DRS capability of ONT has been widely adapted for studies on alternative splicing \u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e,\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e. Native RNA molecules may carry a variety of base modifications, such as m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA and m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC \u003csup\u003e\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e\u003c/sup\u003e, which play roles in the stabilization of RNA transcription and translation. By taking advantage of ONT LR technology, such epi-modifications can be characterized without technique bias. Thus, application of Nanopore-based DRS technology can be expected to greatly advance our understanding of non-coding RNAs and their complex structure, polyadenylation, and modification. In humans, m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA modification has been reported for lincRNAs, \u003cem\u003eMALAT1\u003c/em\u003e, \u003cem\u003eTUG1\u003c/em\u003e, \u003cem\u003eTPT1\u003c/em\u003e, \u003cem\u003eACTB\u003c/em\u003e, and \u003cem\u003eBSG\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e\u003c/sup\u003e. In this study, we found that most lincRNAs in \u003cem\u003eArabidopsis\u003c/em\u003e featured m\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003eA and m\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003eC modification.\u003c/p\u003e \u003cp\u003eIn addition to the above, the long reads from DRS have clear advantages in distinguishing homologs within a complex genome featuring redundant sequences, such as in polyploid species and those sequences from repetitive regions enriched for TEs. ONT DNA sequencing offers exceptionally long reads that permit direct sequencing through regions of the genome previously deemed inaccessible. Although the data presented here are preliminary, we still reveal a clear and systematic difference between groups of molecules at high confidence, namely \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs and non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs.\u003c/p\u003e \u003cp\u003e \u003cb\u003eDe novo\u003c/b\u003e \u003cb\u003eexpression of lincRNAs found in natural populations\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe activation of lincRNAs seen in \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e mutants can also be observed in natural populations, supporting that the non-hypomethylated genome retains the capability to employ its non-coding regions and transcribe RNA from them under given conditions. In the last decade, substantial efforts have been conducted to annotate lincRNAs in multiple species, it still remains a great challenge to identify the functional effects of lincRNA variants in genetic terms \u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e. A fundamental barrier is the identification of regulatory DNA elements on lincRNA loci, which requires a classification system for lincRNA identification and correlation with expression diversity.\u003c/p\u003e \u003cp\u003eFurther classification of such lincRNAs will be facilitated by investigating lincRNA expression polymorphisms in natural populations and identifying \u003cem\u003ede novo\u003c/em\u003e active lincRNAs that have not yet been acted on.\u003c/p\u003e \u003cp\u003e \u003cb\u003ePrediction of\u003c/b\u003e \u003cb\u003eddm1\u003c/b\u003e\u003cb\u003e-activated lincRNAs using machine learning and deep learning\u003c/b\u003e\u003c/p\u003e \u003cp\u003eMachine learning algorithms have been applied to predict genes transcriptionally patterns by training on biochemical features. Here we successfully predict hypomethylation-responsive lincRNAs, which can exhibit more expression diversity. The rank of feature importance within each model helps to reveal possible causes between lincRNA activation and methylation. The dynamics of epigenetic modifications on DNA and RNA are proposed to occur in spontaneous response to environmental and developmental stimuli, in keeping with their roles as constituting the flexible part of the genome. Several DNA methylation loci introduced by TE-associated effects have been reported to serve as critical regulators in crops \u003csup\u003e\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e\u003c/sup\u003e. The integration of machine learning provides an appoach for epigenomics research on lincRNA.\u003c/p\u003e \u003cp\u003eExpression of lincRNAs might occur as a side product of DNA hypomethylation, but hypomethylation-associated lincRNAs have also been reported from crop studies. For example, \u003cem\u003eGhDAN1\u003c/em\u003e in cotton is a non-coding gene reported to regulate drought response by forming a DNA: RNA triplex in leaves. More interestingly, \u003cem\u003eGhDAN1\u003c/em\u003e underwent neofunctionalization in allotetraploid cotton by being released from a hypomethylated chromatin region \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. In addition, a non-coding RNA (\u003cem\u003eAt2g0600\u003c/em\u003e) has been identified as a novel epiallele involved in flowering through simulating selection in a rapidly changing environment and likewise associated with DNA methylation \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Therefore, lincRNAs provide a kind of rapid heritable variation on which natural selection can act. Convolutional neural network models have been successfully applied to predict mRNA transcriptional activation \u003csup\u003e\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e and DNA methylation levels \u003csup\u003e\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e\u003c/sup\u003e based on DNA sequences. Here we successfully predict hypomethylation-related lincRNAs, which can exhibit more expression diversity. Further classification of such lincRNAs will be facilitated by investigating lincRNA expression polymorphisms in natural populations and identifying \u003cem\u003ede novo\u003c/em\u003e active lincRNAs that have not yet been acted on.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eData availability\u003c/h2\u003e \u003cp\u003eAll raw and processed sequencing data generated in this study have been submitted to the NCBI. The high-quality reads and RNA reads from ONT were deposited in the SRA database under accession number BioProject: PRJNA780475.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was financially supported in part by grants from the National Natural Science Foundation of China (NSFC 32000379).\u0026nbsp;We thank Tianzhen Zhang from Zhejiang University for providing insightful suggestions on the populational analysis.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors and Affiliations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eZhejiang Provincial Key Laboratory of Crop Genetic Resources, Institute of Crop Science, Plant Precision Breeding Academy, College of Agriculture and Biotechnology, Zhejiang University, Hangzhou, China\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWanghong Shi, Mengke Zhang, Yupeng Hao, Ke Nie, Xueying Guan \u0026amp; Ting Zhao\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHainan Institute of Zhejiang University, Building 11, Yonyou Industrial Park, Yazhou Bay Science and Technology City, Yazhou District, Sanya, Hainan, 572025, China\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWanghong Shi, Luyao Wang, Mengke Zhang, Yupeng Hao \u0026amp; Ke Nie\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDry land Farming Institute, Hebei Academy of Agricultural and Forestry Sciences, Hengshui, 053000, China\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNa Zhou\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eContributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eT. Z. conceptualized the project. X. G., T. Z., L. W., Y. H. and K. N. performed the experiment. W. S., T. Z., N. Z., M. Z. and X. G. analyzed the data. W. S., T. Z. and X. G. prepared the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cbr\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eLee, H., Zhang, Z., Krause, H.M.: Long Noncoding RNAs and Repetitive Elements: Junk or Intimate Evolutionary Partners? Trends Genet. \u003cb\u003e35\u003c/b\u003e, 892\u0026ndash;902 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalazzo, A.F., Koonin, E.V.: Functional Long Non-coding RNAs Evolve from Junk Transcripts. Cell. \u003cb\u003e183\u003c/b\u003e, 1151\u0026ndash;1161 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWaseem, M., Liu, Y., Xia, R.: Long Non-Coding RNAs, the Dark Matter: An Emerging Regulatory Component in Plants. Int. J. Mol. Sci. \u003cb\u003e22\u003c/b\u003e (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIyer, M.K., et al.: The landscape of long noncoding RNAs in the human transcriptome. Nat. Genet. \u003cb\u003e47\u003c/b\u003e, 199\u0026ndash;208 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndergassen, D., Rinn, J.L.: From genotype to phenotype: genetics of mammalian long non-coding RNAs in vivo. Nat. Rev. Genet. \u003cb\u003e23\u003c/b\u003e, 229\u0026ndash;243 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeo, J.B., Sung, S.: Vernalization-mediated epigenetic silencing by a long intronic noncoding RNA. Science. \u003cb\u003e331\u003c/b\u003e, 76\u0026ndash;79 (2011)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, Y., et al.: Natural temperature fluctuations promote COOLAIR regulation of FLC. Genes Dev. \u003cb\u003e35\u003c/b\u003e, 888\u0026ndash;898 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTao, X., et al.: Neofunctionalization of a Polyploidization-Activated Cotton Long intergenic non-coding RNA DAN1 During Drought Stress Regulation. \u003cem\u003ePlant Physiol\u003c/em\u003e (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, D., et al.: Transposable elements (TEs) contribute to stress-related long intergenic noncoding RNAs in plants. Plant. J. \u003cb\u003e90\u003c/b\u003e, 133\u0026ndash;146 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng, X.M., et al.: Genome-wide analyses reveal the role of noncoding variation in complex traits during rice domestication. Sci. Adv. \u003cb\u003e5\u003c/b\u003e, eaax3619 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMele, M., et al.: Human genomics. The human transcriptome across tissues and individuals. Science. \u003cb\u003e348\u003c/b\u003e, 660\u0026ndash;665 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ede Goede, O.M., et al.: Population-scale tissue transcriptomics maps long non-coding RNAs to complex disease. Cell. \u003cb\u003e184\u003c/b\u003e, 2633\u0026ndash;2648e19 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVakirlis, N., et al.: De novo emergence of adaptive membrane proteins from thymine-rich genomic sequences. Nat. Commun. \u003cb\u003e11\u003c/b\u003e, 781 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKawakatsu, T., et al.: Epigenomic Diversity in a Global Collection of Arabidopsis thaliana Accessions. Cell. \u003cb\u003e166\u003c/b\u003e, 492\u0026ndash;505 (2016)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFurci, L., et al.: Identification and characterisation of hypomethylated DNA loci controlling quantitative resistance in Arabidopsis. Elife \u003cb\u003e8\u003c/b\u003e (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKooke, R., et al.: Epigenetic basis of morphological variation and phenotypic plasticity in Arabidopsis thaliana. Plant. Cell. \u003cb\u003e27\u003c/b\u003e, 337\u0026ndash;348 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe, L., et al.: A naturally occurring epiallele associates with leaf senescence and local climate adaptation in Arabidopsis accessions. Nat. Commun. \u003cb\u003e9\u003c/b\u003e (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGlenfield, C., McLysaght, A.: Pseudogenes Provide Evolutionary Evidence for the Competitive Endogenous RNA Hypothesis. Mol. Biol. Evol. \u003cb\u003e35\u003c/b\u003e, 2886\u0026ndash;2899 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchmid, M.W., et al.: Contribution of epigenetic variation to adaptation in Arabidopsis. Nat. Commun. \u003cb\u003e9\u003c/b\u003e, 4446 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, T., et al.: Population-wide DNA methylation polymorphisms at single-nucleotide resolution in 207 cotton accessions reveal epigenomic contributions to complex traits. Cell. Res. \u003cb\u003e34\u003c/b\u003e, 859\u0026ndash;872 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Z., et al.: lncRNA Epigenetic Landscape Analysis Identifies EPIC1 as an Oncogenic lncRNA that Interacts with MYC and Promotes Cell-Cycle Progression in Cancer. Cancer Cell. \u003cb\u003e33\u003c/b\u003e, 706\u0026ndash;720e9 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, T., et al.: LncRNAs in polyploid cotton interspecific hybrids are derived from transposon neofunctionalization. Genome Biol. \u003cb\u003e19\u003c/b\u003e, 195 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTao, X., et al.: Neofunctionalization of a polyploidization-activated cotton long intergenic non-coding RNA DAN1 during drought stress regulation. Plant. Physiol. \u003cb\u003e186\u003c/b\u003e, 2152\u0026ndash;2168 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, P., et al.: Global Profiling of lncRNAs Expression Responsive to Allopolyploidization in Cucumis. Genes (Basel) 11 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, T., et al.: Absence of CG methylation alters the long noncoding transcriptome landscape in multiple species. FEBS Lett. \u003cb\u003e595\u003c/b\u003e, 1734\u0026ndash;1747 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTan, F., et al.: DDM1 Represses Noncoding RNA Expression and RNA-Directed DNA Methylation in Heterochromatin. Plant. Physiol. \u003cb\u003e177\u003c/b\u003e, 1187\u0026ndash;1197 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStroud, H., Greenberg, M.V., Feng, S., Bernatavichute, Y.V., Jacobsen, S.E.: Comprehensive analysis of silencing mutants reveals complex regulation of the Arabidopsis methylome. \u003cem\u003eCell\u003c/em\u003e 152, 352\u0026thinsp;\u0026ndash;\u0026thinsp;64 (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu, J.M., Johnson, L.M., Jacobsen, S.E., Patel, D.J.: DNA methylation pathways and their crosstalk with histone methylation. Nat. Rev. Mol. Cell Biol. \u003cb\u003e16\u003c/b\u003e, 519\u0026ndash;532 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaw, J.A., Jacobsen, S.E.: Establishing, maintaining and modifying DNA methylation patterns in plants and animals. Nat. Rev. Genet. \u003cb\u003e11\u003c/b\u003e, 204\u0026ndash;220 (2010)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOsakabe, A., et al.: The chromatin remodeler DDM1 prevents transposon mobility through deposition of histone variant H2A.W. Nat. Cell Biol. \u003cb\u003e23\u003c/b\u003e, 391\u0026ndash;U177 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVongs, A., Kakutani, T., Martienssen, R.A., Richards, E.J.: Arabidopsis thaliana DNA methylation mutants. Science. \u003cb\u003e260\u003c/b\u003e, 1926\u0026ndash;1928 (1993)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKankel, M.W., et al.: Arabidopsis MET1 cytosine methyltransferase mutants. Genetics. \u003cb\u003e163\u003c/b\u003e, 1109\u0026ndash;1122 (2003)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLippman, Z., et al.: Role of transposable elements in heterochromatin and epigenetic control. Nature. \u003cb\u003e430\u003c/b\u003e, 471\u0026ndash;476 (2004)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQuadrana, L., Colot, V.: Plant Transgenerational Epigenetics. Annu. Rev. Genet. \u003cb\u003e50\u003c/b\u003e, 467\u0026ndash;491 (2016)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTao, P., et al.: The overexpression of lncRNA MEG3 inhibits cell viability and invasion and promotes apoptosis in ovarian cancer by sponging miR-205-5p. Int. J. Clin. Exp. Pathol. \u003cb\u003e13\u003c/b\u003e, 869\u0026ndash;879 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUlitsky, I., Bartel, D.P.: lincRNAs: Genomics, Evolution, and Mechanisms. Cell. \u003cb\u003e154\u003c/b\u003e, 26\u0026ndash;46 (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKapusta, A., et al.: Transposable elements are major contributors to the origin, diversification, and regulation of vertebrate long noncoding RNAs. PLoS Genet. \u003cb\u003e9\u003c/b\u003e, e1003470 (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBabarinde, I.A., et al.: Transposable element sequence fragments incorporated into coding and noncoding transcripts modulate the transcriptome of human pluripotent stem cells. Nucleic Acids Res. \u003cb\u003e49\u003c/b\u003e, 9132\u0026ndash;9153 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, X., et al.: Full-length transcriptome reconstruction reveals a large diversity of RNA and protein isoforms in rat hippocampus. Nat. Commun. \u003cb\u003e10\u003c/b\u003e, 5009 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEbbert, M.T.W., et al.: Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight. Genome Biol. \u003cb\u003e20\u003c/b\u003e, 97 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGaralde, D.R., et al.: Highly parallel direct RNA sequencing on an array of nanopores. Nat. Methods. \u003cb\u003e15\u003c/b\u003e, 201\u0026ndash;206 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeki, M., et al.: Evaluation and application of RNA-Seq by MinION. DNA Res. \u003cb\u003e26\u003c/b\u003e, 55\u0026ndash;65 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQin, H., et al.: DENA: training an authentic neural network model using Nanopore sequencing data of Arabidopsis transcripts for detection and quantification of N(6)-methyladenosine on RNA. Genome Biol. \u003cb\u003e23\u003c/b\u003e, 25 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWick, R.R., Judd, L.M., Holt, K.E.: Performance of neural network basecalling tools for Oxford Nanopore sequencing. Genome Biol. \u003cb\u003e20\u003c/b\u003e, 129 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie, S., et al.: Applications and potentials of nanopore sequencing in the (epi)genome and (epi)transcriptome era. Innov. (N Y). \u003cb\u003e2\u003c/b\u003e, 100153 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, H.: Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. \u003cb\u003e34\u003c/b\u003e, 3094\u0026ndash;3100 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, H., et al.: The Sequence Alignment/Map format and SAMtools. Bioinformatics. \u003cb\u003e25\u003c/b\u003e, 2078\u0026ndash;2079 (2009)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePertea, M., Kim, D., Pertea, G.M., Leek, J.T., Salzberg, S.L.: Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat. Protoc. \u003cb\u003e11\u003c/b\u003e, 1650\u0026ndash;1667 (2016)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKovaka, S., et al.: Transcriptome assembly from long-read RNA-seq alignments with StringTie2. Genome Biol. \u003cb\u003e20\u003c/b\u003e (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTrapnell, C., et al.: Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat. Protoc. \u003cb\u003e7\u003c/b\u003e, 562\u0026ndash;578 (2012)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTian, L., et al.: (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTrapnell, C., et al.: Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat. Biotechnol. \u003cb\u003e28\u003c/b\u003e, 511\u0026ndash;515 (2010)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKang, Y.J., et al.: CPC2: a fast and accurate coding potential calculator based on sequence intrinsic features. Nucleic Acids Res. \u003cb\u003e45\u003c/b\u003e, W12\u0026ndash;W16 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao, Y., Smyth, G.K., Shi, W.: featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. \u003cem\u003eBioinformatics\u003c/em\u003e 30, 923\u0026thinsp;\u0026ndash;\u0026thinsp;30 (2014)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLove, M.I., Huber, W., Anders, S.: Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. \u003cb\u003e15\u003c/b\u003e, 550 (2014)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorkman, R.E., et al.: Nanopore native RNA sequencing of a human poly(A) transcriptome. Nat. Methods. \u003cb\u003e16\u003c/b\u003e, 1297\u0026ndash;1305 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang, T., et al.: Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. \u003cb\u003e21\u003c/b\u003e, 189 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDanecek, P., et al.: The variant call format and VCFtools. Bioinformatics. \u003cb\u003e27\u003c/b\u003e, 2156\u0026ndash;2158 (2011)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, S., Zhou, Y., Chen, Y., Gu, J.: fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. \u003cb\u003e34\u003c/b\u003e, i884\u0026ndash;i890 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrueger, F., Andrews, S.R.: Bismark: a flexible aligner and methylation caller for Bisulfite-Seq applications. Bioinformatics. \u003cb\u003e27\u003c/b\u003e, 1571\u0026ndash;1572 (2011)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, S., et al.: New insights into Arabidopsis transcriptome complexity revealed by direct sequencing of native RNAs. Nucleic Acids Res. \u003cb\u003e48\u003c/b\u003e, 7700\u0026ndash;7711 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLorenz, D.A., Sathe, S., Einstein, J.M., Yeo, G.W.: Direct RNA sequencing enables m(6)A detection in endogenous transcript isoforms at base-specific resolution. RNA. \u003cb\u003e26\u003c/b\u003e, 19\u0026ndash;28 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao, Y., et al.: Quantitative profiling of N(6)-methyladenosine at single-base resolution in stem-differentiating xylem of Populus trichocarpa using Nanopore direct RNA sequencing. Genome Biol. \u003cb\u003e22\u003c/b\u003e, 22 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim, J., Lee, M., Son, A., Chang, H., Kim, V.N.: mTAIL-seq reveals dynamic poly(A) tail regulation in oocyte-to-embryo development. Genes Dev. \u003cb\u003e30\u003c/b\u003e, 1671\u0026ndash;1682 (2016)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLima, S.A., et al.: Short poly(A) tails are a conserved feature of highly expressed genes. Nat. Struct. Mol. Biol. \u003cb\u003e24\u003c/b\u003e, 1057\u0026ndash;1063 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoman, N.J., Quick, J., Simpson, J.T.: A complete bacterial genome assembled de novo using only nanopore sequencing data. Nat. Methods. \u003cb\u003e12\u003c/b\u003e, 733\u0026ndash;U51 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiang, K., Bartel, D.P.: The molecular basis of coupling between poly(A)-tail length and translational efficiency. Elife \u003cb\u003e10\u003c/b\u003e (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmeur, A., Kloosterman, W.P., Hestand, M.S.: Single-Molecule Sequencing: Towards Clinical Applications. Trends Biotechnol. \u003cb\u003e37\u003c/b\u003e, 72\u0026ndash;85 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, N., et al.: Probing N6-methyladenosine RNA modification status at single nucleotide resolution in mRNA and long noncoding RNA. RNA. \u003cb\u003e19\u003c/b\u003e, 1848\u0026ndash;1856 (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuttenhofer, A., Schattner, P., Polacek, N.: Non-coding RNAs: hope or hype? Trends Genet. \u003cb\u003e21\u003c/b\u003e, 289\u0026ndash;297 (2005)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu, G., et al.: Evolutionary and functional genomics of DNA methylation in maize domestication and improvement. Nat. Commun. \u003cb\u003e11\u003c/b\u003e, 5539 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWashburn, J.D., et al.: Evolutionarily informed deep learning methods for predicting relative transcript abundance from DNA sequence. Proc. Natl. Acad. Sci. U S A. \u003cb\u003e116\u003c/b\u003e, 5542\u0026ndash;5549 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeng, H., Gifford, D.K.: Predicting the impact of non-coding variants on DNA methylation. Nucleic Acids Res. \u003cb\u003e45\u003c/b\u003e, e99 (2017)\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Direct RNA-sequencing, non-coding RNA, ddm1, met1, Arabidopsis, hypomethylation, machine learning","lastPublishedDoi":"10.21203/rs.3.rs-6399991/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6399991/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLong non-coding RNAs are regulatory RNAs with significantly high expression diversity within populations. In both plant and animal genomes, expression of long intergenic non-coding RNAs (lincRNAs) is associated with reduced genomic DNA methylation in non-coding regions. Whether such newly-activated lincRNAs are widespread in natural populations and have predictable patterns associated with natural selection has not been investigated. Here, we employed Oxford Nanopore Technology Direct RNA and DNA sequencing (ONT DRS and DDS) in DNA methylation-deficient \u003cem\u003eArabidopsis\u003c/em\u003e mutants \u003cem\u003eddm1\u003c/em\u003e and \u003cem\u003emet1\u003c/em\u003e to generate 41\u0026nbsp;million high-quality long RNA reads. In total, 340 lincRNAs were found to be activated under defective DNA methylation, while 209 were constitutively expressed in both mutants and WT. Expression of \u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs was negatively correlated with DNA methylation levels. DNA-hypomethylation-activated lincRNAs can be detected in natural populations at low frequency and with high expression diversity. The \u003cem\u003eddm1\u003c/em\u003e-activated and non-\u003cem\u003eddm1\u003c/em\u003e-activated lincRNAs can be distinguished by Random Forest algorithm and Deep convolutional neural network with up to 70% and 91% accuracy. The integrated results suggest that dynamics of DNA methylation in non-coding regions are associated with non-coding RNAs in a predictable fashion with potential new function. This work advances a framework for leveraging epigenomic signatures in non-coding RNA discovery.\u003c/p\u003e","manuscriptTitle":"Nanopore Direct RNA-Seq Reveals Widespread and Predictable Non-Coding Transcriptional Variations in DNA Methylation-deficient Arabidopsis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-07 17:28:41","doi":"10.21203/rs.3.rs-6399991/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"communications-biology","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commsbio","sideBox":"Learn more about [Communications Biology](http://www.nature.com/commsbio/)","snPcode":"","submissionUrl":"","title":"Communications Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"48a24161-7642-40c9-a792-4b64c50feaa3","owner":[],"postedDate":"May 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":47965219,"name":"Biological sciences/Genetics/Epigenetics/DNA methylation"},{"id":47965220,"name":"Biological sciences/Genetics/Genomics/Transcriptomics"}],"tags":[],"updatedAt":"2026-02-07T08:11:26+00:00","versionOfRecord":{"articleIdentity":"rs-6399991","link":"https://doi.org/10.1038/s42003-025-09454-1","journal":{"identity":"communications-biology","isVorOnly":false,"title":"Communications Biology"},"publishedOn":"2025-12-31 05:00:00","publishedOnDateReadable":"December 31st, 2025"},"versionCreatedAt":"2025-05-07 17:28:41","video":"","vorDoi":"10.1038/s42003-025-09454-1","vorDoiUrl":"https://doi.org/10.1038/s42003-025-09454-1","workflowStages":[]},"version":"v1","identity":"rs-6399991","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6399991","identity":"rs-6399991","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.