Building a Model for Predicting Target Gene Expression in Rice T-DNA Insertional Mutants by Machine Learning Approaches

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background T-DNA activation-tagging technology is widely used to enhance flanking gene expression near the site of insertion for functional genomics research in rice. However, whether the expression of a gene of interest is enhanced must be validated experimentally. Results In this study, we built a model to predict gene expression in T-DNA mutants by machine learning approaches, thereby improving the efficiency of screening for activated genes. We gathered experimental consisting of gene expression data in T-DNA mutants and captured the PROMOTER and MIDDLE sequences for encoding. In first-layer models, SVM models were constructed with nine features consisting of information about biological function and local and global sequences. Feature-encoding based on the PROMOTER sequence was weighted by logistic regression. The second-layer models integrated 16 first-layer models with feature selection and the algorithm, which were selected from nine feature selection methods and 65 classified methods, respectively. The accuracy of the final two-layer machine learning model, referred to as was 99.3% based on five-fold cross-validation, and 85.6% based on independent-testing. Conclusion We discovered that the information within the local sequence had a greater contribution than the global sequence with respect to classification had a good predictive ability for target genes within 20 from the 35S enhancer. Based on the analysis of significant sequences, the G-box regulatory sequence may also play an important role in the mechanism of activation of the 35S enhancer.
Full text 16,948 characters · extracted from preprint-html · click to expand
Building a Model for Predicting Target Gene Expression in Rice T-DNA Insertional Mutants by Machine Learning Approaches | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Building a Model for Predicting Target Gene Expression in Rice T-DNA Insertional Mutants by Machine Learning Approaches Chi-Chou Liao, Liang-Jwu Chen, Shuen-Fang Lo, Chi-Wei Chen, Jia-Jyun Chen, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-20492/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background T-DNA activation-tagging technology is widely used to enhance flanking gene expression near the site of insertion for functional genomics research in rice. However, whether the expression of a gene of interest is enhanced must be validated experimentally. Results In this study, we built a model to predict gene expression in T-DNA mutants by machine learning approaches, thereby improving the efficiency of screening for activated genes. We gathered experimental consisting of gene expression data in T-DNA mutants and captured the PROMOTER and MIDDLE sequences for encoding. In first-layer models, SVM models were constructed with nine features consisting of information about biological function and local and global sequences. Feature-encoding based on the PROMOTER sequence was weighted by logistic regression. The second-layer models integrated 16 first-layer models with feature selection and the algorithm, which were selected from nine feature selection methods and 65 classified methods, respectively. The accuracy of the final two-layer machine learning model, referred to as was 99.3% based on five-fold cross-validation, and 85.6% based on independent-testing. Conclusion We discovered that the information within the local sequence had a greater contribution than the global sequence with respect to classification had a good predictive ability for target genes within 20 from the 35S enhancer. Based on the analysis of significant sequences, the G-box regulatory sequence may also play an important role in the mechanism of activation of the 35S enhancer. Bioinformatics Rice CaMV 35S enhancer T-DNA activation tagging Gene expression Machine learning Figures Figure 1 Figure 2 Figure 3 Figure 4 Full Text Supplementary Files TIMgoSupplementary.docx SupplementaryFiles.rar Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-20492","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":497086,"identity":"d9b42fee-cd3c-45cc-8698-8521ed1d9db8","order_by":1,"name":"Chi-Chou Liao","email":"","orcid":"","institution":"National Chung Hsing University College of Life Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Chi-Chou","middleName":"","lastName":"Liao","suffix":""},{"id":497087,"identity":"a9727558-b586-4ea2-9381-902113ea938f","order_by":2,"name":"Liang-Jwu Chen","email":"","orcid":"","institution":"National Chung Hsing University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Liang-Jwu","middleName":"","lastName":"Chen","suffix":""},{"id":497088,"identity":"b2f2e37d-c253-4e7b-881a-5d79445bde8b","order_by":3,"name":"Shuen-Fang Lo","email":"","orcid":"","institution":"Academy of Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shuen-Fang","middleName":"","lastName":"Lo","suffix":""},{"id":497089,"identity":"309bd4e0-b1a6-421b-9dcd-4dd5eec174fe","order_by":4,"name":"Chi-Wei Chen","email":"","orcid":"","institution":"National Chung Hsing University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Chi-Wei","middleName":"","lastName":"Chen","suffix":""},{"id":497090,"identity":"ad88207a-4ee9-4369-8bf5-af83c1bd1d8a","order_by":5,"name":"Jia-Jyun Chen","email":"","orcid":"","institution":"National Chung Hsing University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jia-Jyun","middleName":"","lastName":"Chen","suffix":""},{"id":497091,"identity":"5be59db1-c586-42ef-926c-4ac6b243d961","order_by":6,"name":"Yen-Wei Chu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAtElEQVRIiWNgGAWjYHACNgaGCgYGAxCTh3gtZxBaJIjTwthGihaDG8nPHvPOuyNvLpHA+OBtG0OdwQGCWtLMjXm3PTPcOSOB2XBuG4MEYS23c9ikebcdTjC4kQBkALWYEadlDlgL+28StDRAbGEmSovk/WdmknOOHTbccOZhs+SccxKS+wlp4Ttz+JnEm5rD8gbHkw9+eFNmwy/ZQECLAsJMRpBaImJSnpCZo2AUjIJRMAoYAKrPPyQVavaYAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-5525-4011","institution":"National Chung Hsing University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Yen-Wei","middleName":"","lastName":"Chu","suffix":""}],"badges":[],"createdAt":"2020-03-31 11:45:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-20492/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-20492/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":950889,"identity":"01c86122-631e-4f84-a8ca-829968dbc64d","added_by":"auto","created_at":"2020-04-22 19:43:15","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":577592,"visible":true,"origin":"","legend":"Flowchart of the TIMgo predictive system. TIMgo was built in two-layer model, primary module included nine feature-encoding and secondary modules was integrated nine results of primary modules. The red dash line indicates the system core architecture.","description":"","filename":"figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/figure1.png"},{"id":950892,"identity":"3d54e08e-63fd-4f05-9c26-b4c18e2ea3ab","added_by":"auto","created_at":"2020-04-22 19:43:16","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":212143,"visible":true,"origin":"","legend":"Correlation between distance and gene activation. The data were sorted by the distance between the 35S enhancer and the TLS, and the ratio of Ac to NAc genes in each group was calculated. The x axis is the distance from the 35S enhancer to the TLS of a target gene; the y axis is the proportion of Ac and NAc genes in each group.","description":"","filename":"figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/figure2.png"},{"id":950893,"identity":"3ba92445-7087-42a4-ac0e-c114e34bd2ef","added_by":"auto","created_at":"2020-04-22 19:43:16","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":634034,"visible":true,"origin":"","legend":"Accuracy trend in the second-layer feature selection. The x-axis represents how many features the models used, and y-axis represents the accuracy of model had been built with some of features. In this study, nine feature selection methods were used.","description":"","filename":"figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/figure3.png"},{"id":950894,"identity":"28c27853-c4c3-464c-b4fc-df5379c0ec21","added_by":"auto","created_at":"2020-04-22 19:43:16","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":231069,"visible":true,"origin":"","legend":"Accuracy trend of TIMgo for cross-validation and independent-testing of data within different distances. Train represents the Acc from 5-fold cross validation with D299; Test represents the Acc from independent testing with D153. The x axis indicates each distance interval, and the y axis indicates the predictive accuracy.","description":"","filename":"figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/figure4.png"},{"id":1116413,"identity":"2bdee2b2-d374-453e-87b1-485705895ec5","added_by":"auto","created_at":"2020-05-17 21:39:38","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":559235,"visible":true,"origin":"","legend":"","description":"","filename":"TIMgoManuScript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/TIMgoManuScript.pdf"},{"id":950895,"identity":"b42a3590-bf95-47bb-8948-cd2f89113945","added_by":"auto","created_at":"2020-04-22 19:43:26","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":616010,"visible":true,"origin":"","legend":"","description":"","filename":"TIMgoManuScript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/Manuscript.pdf"},{"id":950890,"identity":"1ee81cd2-024b-44ba-b209-fcd006140c7c","added_by":"auto","created_at":"2020-04-22 19:43:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":559235,"visible":true,"origin":"","legend":"","description":"","filename":"TIMgoManuScript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/TIMgoManuScript.pdf"},{"id":13499717,"identity":"7b977b76-3b5e-4370-9755-b24e16ebbced","added_by":"auto","created_at":"2021-09-16 23:03:26","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":767594,"visible":true,"origin":"","legend":"","description":"","filename":"TIMgoManuScript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1_covered.pdf"},{"id":950891,"identity":"dcb84588-80dd-4b53-86d7-1e61c16de7ed","added_by":"auto","created_at":"2020-04-22 19:43:16","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":182638,"visible":true,"origin":"","legend":"","description":"","filename":"TIMgoSupplementary.docx","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/TIMgoSupplementary.docx"},{"id":950888,"identity":"fdfd86f4-ba90-4c23-9242-68cb4e2267e4","added_by":"auto","created_at":"2020-04-22 19:43:15","extension":"rar","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":224366,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFiles.rar","url":"https://assets-eu.researchsquare.com/files/rs-20492/v1/SupplementaryFiles.rar"}],"financialInterests":"","formattedTitle":"Building a Model for Predicting Target Gene Expression in Rice T-DNA Insertional Mutants by Machine Learning Approaches","fulltext":[{"header":"Full Text","content":"\u003cp\u003eThis preprint is available for \u003ca href='/article/rs-20492/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Rice, CaMV 35S enhancer, T-DNA activation tagging, Gene expression, Machine learning","lastPublishedDoi":"10.21203/rs.3.rs-20492/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-20492/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eT-DNA activation-tagging technology is widely used to enhance flanking gene expression near the site of insertion for functional genomics research in rice. However, whether the expression of a gene of interest is enhanced must be validated experimentally.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eIn this study, we built a model to predict gene expression in T-DNA mutants by machine learning approaches, thereby improving the efficiency of screening for activated genes. We gathered experimental consisting of gene expression data in T-DNA mutants and captured the PROMOTER and MIDDLE sequences for encoding. In first-layer models, SVM models were constructed with nine features consisting of information about biological function and local and global sequences. Feature-encoding based on the PROMOTER sequence was weighted by logistic regression. The second-layer models integrated 16 first-layer models with feature selection and the algorithm, which were selected from nine feature selection methods and 65 classified methods, respectively. The accuracy of the final two-layer machine learning model, referred to as was 99.3% based on five-fold cross-validation, and 85.6% based on independent-testing.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eWe discovered that the information within the local sequence had a greater contribution than the global sequence with respect to classification had a good predictive ability for target genes within 20 from the 35S enhancer. Based on the analysis of significant sequences, the G-box regulatory sequence may also play an important role in the mechanism of activation of the 35S enhancer.\u003c/p\u003e","manuscriptTitle":"Building a Model for Predicting Target Gene Expression in Rice T-DNA Insertional Mutants by Machine Learning Approaches","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-04-22 19:43:13","doi":"10.21203/rs.3.rs-20492/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dc34635f-6031-4d0b-9026-cb71d8289bb7","owner":[],"postedDate":"April 22nd, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":86550,"name":"Bioinformatics"}],"tags":[],"updatedAt":"2020-05-17T21:39:37+00:00","versionOfRecord":[],"versionCreatedAt":"2020-04-22 19:43:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-20492","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-20492","identity":"rs-20492","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00