Unveiling insights into drought-tolerant responses in soybean: a data-driven pipeline for feature engineering in biomarker discovery | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Unveiling insights into drought-tolerant responses in soybean: a data-driven pipeline for feature engineering in biomarker discovery Pei-Hsiu Kao, Supaporn Baiya, Chong-Wei Lee, Chia-Wen Tseng, Shu-Yun Chen, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4791767/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 19 Mar, 2025 Read the published version in Journal of Big Data → Version 1 posted 13 You are reading this latest preprint version Abstract Soybean [ Glycine max (L.) Merr.] is an important global food crop but is highly vulnerable to environmental changes, particularly drought. Conventional strategies of biomarkers discovery for developing drought-tolerant varieties are resource-intensive, inefficient, and without comprehensive insight. The current study proposed a novel data-driven pipeline for feature engineering through integrating diverse genetic data from multidisciplinary research on cloud-based sources in identification of key drought-tolerant genes (DTgenes) in soybean. Our pipeline involved data extraction, transformation, loading, and systematic integration of both omics and non-omics data. Feature prioritization was performed for feature selection to uncover the important biomarkers from feature pool (candidate gene pool), and key DTgenes were identified through integrative systems biology strategies based on the biomarker candidates. The validation of identified key DTgenes was carried out by both computational and molecular experiments approach. Both approaches demonstrated the credibility and potential of the key DTgenes for conferring drought tolerance response in soybean. This data-driven pipeline for feature engineering approach enhances the efficiency and accuracy of biomarkers discovery for further breeding program, which uncover a robust key DTgenes candidate which contribute to drought tolerance improvement in soybean and show potential of applicability for other crop species. soybean drought-tolerant responses data-driven systems biology multidisciplinary biomarkers discovery Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Introduction Soybean ( Glycine max L.) is an important global food crop owing to its high protein and oil content, but it faces high susceptibility to environmental fluctuations, particularly water availability [1]. Extreme weather caused by the climate change incurs the escalating of drought events, which poses a substantial threat to soybean cultivation and further reduce agricultural productivity [2, 3]. Drought stress primarily results from water scarcity and inadequate soil moisture supply, inducing various adverse effects on crop physiology [4]. Persistent water deficit during the vegetative growth forces early transition to reproductive stage, leading to substantial yield losses [5]. Furthermore, drought stress during early reproductive or grain-filling stages incurs pollen sterility and impaired grain set, exacerbating economic losses [6, 7]. Conventional breeding is time-consuming practice for developing new crop varieties [8]. The time constraints, inherent genetic stochasticity, environmental fluctuations, and substantial resource investments set formidable obstacles to breeding program [9, 10]. The rapid pace of climate change frequently surpasses the timeline for developing new and resilient varieties, exacerbating agricultural losses. To overcome these issues, there is a growing demand for more rapid, accurate, and cost-efficient methodologies in identifying key resistant genes for tolerant varieties development. A data-driven approach emerges as a paradigm-shifting solution by applying extensive biological datasets. The integration of multidisciplinary research data and systems biology approaches provides novel insights into crop improvements by the use of comprehensive information, facilitating the breeding progress [11]. The current understanding of drought tolerance in soybean still remains limited. Although some advanced high-throughput techniques have been developed and applied to biomarkers discovery for drought improvement, it still remains challenges due to the complex interplay of multiple biological processes in drought-tolerant response. Drought-tolerant response cannot be explored and interpreted relying solely on a single study or unilateral molecular approach, such as genome-wide association study (GWAS) or RNA-seq, which may not capture enough information. For instance, current GWAS research can only account for approximately 30% of the phenotypic variance observed between drought-tolerant lines and commercial lines. This suggests that the contribution of each single nucleotide polymorphism (SNP) to the variance explainability is exceedingly low [12, 13]. Other non-omics or non-whole genome studies also face challenges for biomarkers discovery. The variations in materials, methods, and environmental conditions applied in individual study incurs issues in standardization and reproducibility of the experiment outcome [14, 15]. To expand our understanding to the remaining phenotypic variance, there is an urgent need for integrative and interdisciplinary approaches. Integrative approaches, such as combining omics and non-omics (OnO) data, have shown promise in uncovering the genetic mechanisms underlying complex traits in various organisms [16], indicating their potential applicability in soybean drought tolerance research. The raising of high-throughput techniques, propelled by advancements in computational capabilities and software tools, creates a new era for biology [17]. This paradigm shift poses challenges in identification of important information from the soaring volume of biological data. The integration of multi-dimensional biological data presents an opportunity to gain comprehensive insights into the molecular mechanisms underlying complex traits in soybean. This integration process, however, encounters certain challenges, primarily due to the heterogeneity of data types [18], requiring specialized tools for effective handling. Structured omics data offers standardized formats and well-defined variables, facilitating analysis and interpretation. On the other hand, unstructured data sources, including textual information from scientific literature and phenotypic data from field trials, lack a predefined structure, posing challenges in data organization and integration. Furthermore, variability in data quality and completeness across different datasets can introduce biases and noise [19]. Despite these challenges, the integration of multi-dimensional biological data holds immense promise for advancing soybean breeding and genomics research by providing a holistic understanding of the genetic factors governing agronomic traits [20–22]. Data mining plays a critical role in the extraction of meaningful information from large-scale datasets [23]. A large challenge pertains to the abundance of non-structured data, which constitute over 80% of valuable data pool [24, 25]. Text mining stands out as particularly significant for wanted information extracted from diverse data types, including structured, semi-structured, quasi-structured, and unstructured formats. Structured data, characterized by predefined fields and formats, including relational databases and spreadsheets, whereas unstructured data, such as images, text files, audio files, and PDFs, lacks a rigid organizational schema. Semi-structured data, represented by formats like comma-separated value (CSV), extensible markup language (XML), and hypertext markup language (HTML), occupies an intermediate position between structured and unstructured data, with part of organizational properties but lacking uniformity in database structure [26]. Quasi-structured data, although less frequently discussed, the irregular formats requires specialized tools for efficient extraction and processing [27]. A comprehensive understanding of diverse data types is indispensable in bioinformatics, facilitating the effective utilization of the wide spectrum of data essential for our research endeavors. Data extraction-transformation-loading (ETL) serves as a fundamental process for the cleansing, customization, and reformatting of data within the realm of data warehousing [28]. By using data-driven approaches, the ETL process enables the systematic extraction of relevant information from diverse biological databases, including OnO data. Initially, raw data is extracted from various sources databases, along with textual information from scientific literature and field trials. Subsequently, data cleaning techniques are applied to remove inconsistencies, errors, and irrelevant data [29, 30], ensuring the quality and integrity of the dataset [30]. This systematic approach is crucial for maintaining data quality, particularly considering the non-standardized nature of biological data [31]. Data filtering procedures are then implemented to focus on specific features relevant to drought tolerance traits in soybean. Additionally, data mapping techniques facilitate the integration and alignment of heterogeneous data sources, allowing for the correlation of genetic information with phenotypic traits. Following this, data sorting algorithms are employed to organize the dataset according to predefined criteria in data repository or warehouse, facilitating efficient analysis and interpretation. Finally, the processed data is loaded into computational tools and databases for further analysis and feature selection. Systems biology, including pathway enrichment analysis, network analysis, and topology properties, represents a powerful tool to understand cellular and organismal functions holistically at a systemic level [32]. Through the integration of diverse OnO data and the application of mathematical models, systems biology offers a comprehensive approach to analyzing and interpreting heterogeneous datasets, thereby enhancing our understanding of plant physiology and biological systems [33, 34]. By aggregating information from multiple genes into annotated pathways, such as Gene Ontology (GO) terms, the joint effects of numerous genes can be assessed to identify enriched biological pathways [35] that are modulated in response to a specific trait. Network analysis allows for the construction of interaction networks among genes, proteins, and other molecular entities involved in the trait, revealing the intricate interactions between genes and pathways [36]. Furthermore, topology-based analysis, such as degree, centrality measures, connectivity, and clustering coefficients, can effectively model biological phenomena by incorporating pathways to identify key genes within the network that play crucial roles in the trait. Together, these approaches collectively offer a systems-level perspective of the molecular landscape associated with a specific trait, facilitating the prioritization of key genes and pathways for further investigation and targeted breeding strategies. In the current study, we introduced a novel data-driven pipeline for key genes discovery, which aimed to overcome the limitations of conventional breeding practices for crop improvement. This innovative strategy harnesses the flexibility of data-driven methodologies, combining diverse data from multi-disciplinary research. It provides a more comprehensive and efficient approach to address the complex genotype-phenotype associations to support decision-making in breeding programs. The proposed data-driven pipeline not only offers a promising solution for identifying key genes for crop improvement but also serves as a protocol for future research and applications across various fields. Materials and methods A data-driven pipeline for feature engineering We devised a data-driven pipeline for feature engineering, termed DPFE, and aimed to identify key genes associated with soybean drought-tolerant responses. The DPFE system comprises six layers: a data source layer involving cloud-based data mining and wrangling, a data processing layer encompassing data extraction, transformation, loading, and harmonization to retrieve wanted information from diverse data types and enhance data quality, a data integration layer where processed data is integrated into training and validation gene sets, a data prioritization layer utilizing feature selection to identify biomarkers, a data systems layer employing pathway enrichment, gene networks, and topological analysis for key DTgenes discovery, and an evaluation layer assessing key DTgenes through computational methods and molecular experiments. The comprehensive architecture of DPFE is depicted in Fig. 1 . Data mining and wrangling ( data source layer ) To comprehensively elucidate the genetic mechanisms underlying drought-tolerant responses in soybean, we initiated an extensive data acquisition process across multiple sources, including data warehouses, data lakes, and cloud-based experimental repositories. Relevant datasets were systematically retrieved using targeted search terms on platforms such as NCBI, focusing specifically on drought tolerance-related data. This included genomics, transcriptomics, proteomics, metabolomics, and epigenomics data pertinent to soybean under drought conditions. Following data mining, rigorous data wrangling protocols were implemented to ensure data quality and relevance. The protocols involved data curation and categorization. We implemented strict curation protocols to filter and retain only data directly relevant to drought-tolerant responses. This involved retaining datasets from field trials with controlled soil water content and experiments simulating drought conditions using varying levels of polyethylene glycol. These protocols were crucial to not only align the data to our research objective by laying the foundation onto the data of biological relevance and integrity, but also minimize the noise and maximize the signal-to-noise ratio. To effectively manage the diversity of data, we categorized datasets based on their structural and formatting characteristics. This included distinguishing between structured data (e.g., spreadsheets) and non-structured data. Non-structured data incorporated semi-structured data (e.g., HTML, CSVs, PDFs with extractable tables), quasi-structured data (e.g., PDFs without extractable tables), and unstructured data (e.g., images). This categorization was pivotal for organizing the training gene set and guiding subsequent data processing and analysis methods tailored to each data type. Data ETL process ( data processing layer ) We devised a comprehensive data-driven ETL process combined with robust feature engineering techniques to identify key genes associated with soybean drought-tolerant responses. The ETL process is a powerful procedure to retrieve wanted information from multiple data sources, maximizing the utility of available data for more accurate and insightful discoveries. This process is vital for organizing and preparing data, ensuring its quality and relevance to inform our analysis. Extraction We employed Python v3.11.4 for data extraction across various sources including structured, semi-structured, quasi-structured, and unstructured data. Structured data was processed with Pandas 2.0.3 for spreadsheets. Semi-structured data was accessed through BeautifulSoup v4.21.1, Pandas 2.0.3, and Tabula v1.0.5 for web scraping from HTML, CSV files, and PDF files with extractable tables, respectively. Quasi-structured data extraction required Tabula v1.0.5 and extensive manual curation for non-extractable PDF tables. Unstructured data, particularly images, underwent manual annotation to identify relevant features. Semi-structured and quasi-structured data required customized scripts for processing, whereas unstructured data involved more intensive manual curation due to its complexity. Transformation Following extraction, data underwent a rigorous transformation process. This included cleaning, filtering, and standardizing to eliminate irrelevant or erroneous data. Structured data was reformatted and categorized based on its relevance to soybean drought-tolerant response. Non-structured data was converted into structured formats, enabling seamless integration into our data platforms. We meticulously categorized the transformed data into appropriate OnO data platforms to maintain clarity. Loading Data loading involved systematic integration of the transformed and cleaned data into a structured data repository. This repository serves as a centralized storage system, allowing for efficient data retrieval and analysis. Ensuring data integrity and consistency at this stage is critical for subsequent analysis. To maintain consistency, we introduced standardized columns to align the new data with the existing data structure. The standardized columns include data type (structured, semi-structured, quasi-structured, unstructured), data format (spreadsheet, HTML, CSV, PDF, image), OnO (omics, non-omics), system component level (genome, transcriptome, proteome, metabolome, epigenome), feature (SNP, gene, SSR, QTL), chromosome, start genomic position, end genomic position, genome version, drought duration, statistical significance value (p-value), measure of linkage strength (logarithm of the odds, LOD), measure of differential expression (FC), computed value representing feature importance (score), developmental stage of the plant (growth stage), and source or reference of the data (reference). These standardized columns ensure that all data entries adhere to a uniform format, facilitating easier data manipulation and analysis. A final quality assurance step is conducted to ensure that the data in the repository meets the required standards for completeness, accuracy, and relevance. Any issues identified during this step are addressed before the data is made available for analysis. Multiple feature integration ( Data integration layer ) The integration of multiple data features is fundamental to capture the complexity of biological systems. By combining different data features, we can achieve a more comprehensive understanding of the genotype-phenotype associations that underpin drought-tolerant responses in soybean. In this study, we addressed the challenge of integration of heterogeneous data from various omics (OnO) level of the system components (denoted as data platforms), which include diverse data types and originate from different molecular methodologies, plant materials, and environmental conditions. Recognizing the inherent risk of FPs in molecular biology experiments, our primary objective was to minimize this risk in the discovery of key DTgenes. Our hypothesis was that biomarkers genuinely associated with drought-tolerant responses would exhibit consistent patterns across various OnO data platforms, while irrelevant biomarkers, such as FPs, would not. Those FPs would neither achieve high scores across data platforms nor reach top-ranking positions, making them less likely to be selected as DTgenes candidates (Fig. 2 ). To tackle this, we employed conceptual integration along with integrative feature selection strategy based on the principle that the whole outperforms the sum of its parts [37]. This approach involved harmonizing and normalizing different data points across multiple OnO data platforms using a scoring scheme. Here, we partitioned the entire dataset into the training gene set and the validation gene set. Multifaceted feature integration (constructing the training gene set) We developed a comprehensive approach for assessing drought-tolerant responses in soybean through the integration of multiple features. Specifically, we established a drought-tolerant response score ( \(\:{S}_{j}^{DR}\) ) for each feature in the training gene set calculated as the product of the sum of genetic scores across multiple OnO data platforms ( \(\:{S}_{.j}^{G}\) ) and their respective importance index ( \(\:{II}_{j}\) ) to represent the magnitude of association between the feature and the trait. The genetic score \(\:{S}_{ij}^{G}\) was determined using a score transformation function, denoted as \(\:{h}^{ij}(\bullet\:)\) , applied to synthesize a single or several data points from each OnO data platform (indexed by i ) for a specific feature (indexed by j ). We summed \(\:{S}_{ij}^{G}\) across all OnO data platforms and subsequently rescaled it to a standardized range between 0 and 6. To evaluate the significance of each feature in relation to the trait, we introduced the important index ( \(\:{II}_{j}\) ). Initially, \(\:{II}_{j}\) was assigned a value of 1.0 to represent a neutral starting point for each feature and adjusted based on its performance threshold ( p -value 1.5, or FC > 1.5) within individual OnO data platform. Specifically, \(\:{II}_{j}\) increased by 0.2 if the feature met the threshold, and decreased by 0.1 if it did not. Features with indeterminate performance maintained their initial \(\:{II}_{j}\) value. This approach allows us to systematically quantify the importance of each feature in contributing to soybean drought tolerance. By integrating data from multiple platforms and considering their respective significance, we aim to identify key genes associated with drought-tolerant responses in soybean. In our methodology, we utilized diverse scoring schemes and data normalization techniques specific to each OnO data platform to calculate the genetic score \(\:{S}_{.j}^{G}\) for individual features (indexed by j ). Given the heterogeneous nature of the data, including various types of numerical genetic information such as p -values (for statistical significance in association mapping data), LOD scores (for assessing the strength of genetic linkage in linkage or association mapping data), and FC values (for evaluating the magnitude of expression differences in expression data), derived from molecular characteristics within each OnO data platform, it was essential to employ platform-specific procedures for score calculation. To ensure comparability and consistency across platforms, we implemented appropriate data normalization techniques tailored to the characteristics of each dataset. These normalization procedures aimed to account for inherent variations in data distributions and experimental conditions, thereby facilitating meaningful comparisons and integrations across multiple platforms. We employed various transformation functions to combine multiple \(\:\pi\:\) values obtained from \(\:m\) independent hypothesis tests \(\:({H}_{0}:{\pi\:}_{s}\sim\text{U}\left[\text{0,1}\right],\:s=1,\cdots\:,m)\) into a single test statistic for the j -th feature within the i -th OnO data platform. Specifically, we used the following transformation functions: \(\:{h}_{P}^{ij}\left(\bullet\:\right)=-2\sum\:_{s=1}^{m}\text{ln}({\pi\:}_{s})\) for combining p -values, which is distributed as \(\:{x}_{2m}^{2}\) under the null hypothesis. \(\:{h}_{NP}^{ij}\left(\bullet\:\right)=\text{m}\text{i}\text{n}\left\{{\pi\:}_{1},{\pi\:}_{2},\cdots\:,{\pi\:}_{m}\right\}\) for combining non- p -value genetic information, which follows a \(\:\text{B}\text{e}\text{t}\text{a}(1,m)\) distribution under the null hypothesis. To process non- p -value data (such as LOD scores, FC values, degrees, and pathway scores), we employed a reciprocal of the logarithmically transformed absolute value of various genetic information values to convert these values into \(\:\pi\:\) values. This transformation is mathematically represented as \(\:{10}^{-\text{log}\left(\right|.\left|\right)}\) , which ensures that the values are normalized on a comparable scale. The genetic score \(\:{S}_{.j}^{G}\) was then calculated using the formula \(\:{-\text{l}\text{o}\text{g}\{h}^{ij}(\bullet\:)\}\) , where \(\:{h}^{ij}\left(\bullet\:\right)=\text{m}\text{i}\text{n}\left\{{h}_{P}^{ij}\left(\bullet\:\right),{\:h}_{NP}^{ij}\left(\bullet\:\right)\:\right\}\) . A higher value of \(\:{S}_{.j}^{G}\) suggested a stronger potential \(\:{S}_{.j}^{G}\) connection with drought-tolerant responses in soybean. These transformation and normalization procedures were designed to facilitate meaningful comparisons and integrations across multiple platforms, ensuring that the diverse genetic information could be accurately and consistently interpreted. We designated this integrated dataset as the training gene set. Validation gene set We selected two omics datasets, which were excluded from the training gene set, as the validation gene set. The first selected transcriptomics dataset contains 104,022 gene expression profiles, encompassing three different drought treatments (6h, 12h, and 24h) conducted on R2 soybean seedlings [38]. The second selected transcriptomics dataset contains 112,094 gene expression profiles, encompassing three different drought treatments (5 days and 7 days) conducted on V1 soybean seedlings [39]. Omics swap The ‘omics swap’ procedure is a validation technique designed to ensure the robustness of selected biomarkers by testing their consistency across different data subsets. This involves alternating the roles of a training gene set and a validation set within the integrated feature analysis framework. In this study, we selected one genomics dataset as the new validation set. This dataset contains 139,428 SNPs, mapped to 49,662 genes, with soybean seedlings subjected to 3–4 weeks and 6–7 weeks of drought treatment [12]. The goal is to verify that the identified features maintain their significance and predictive power irrespective of the data subset used. Feature selection and prioritization ( data prioritization layer ) Feature prioritization Feature selection and prioritization is a crucial step in identifying key biomarkers associated with drought tolerance responses in soybean. To achieve this, we employed the NPRF algorithm on the training gene set, which is capable of handling high-dimensional data and proficient at selecting a subset of features strongly linked to drought-tolerant responses. The NPRF algorithm constructs decision trees using bootstrap resampling and random permutation. This approach enables the calculation of the frequency of each biomarker by comparing the average scores of bootstrap samples with permutation samples through the Wilcoxon rank-sum test. Specifically, we conducted 1,000 bootstrap iterations and performed 10 rounds of random permutation within each bootstrap iteration to compute the gene frequencies. An appropriate threshold ( k ) was then determined, corresponding to a dramatical drop in gene frequency, to identify the most significant features. The top k features with the highest drought-tolerant response score ( \(\:{S}_{j}^{DR})\) were selected as our initial set of candidate biomarkers. To ensure a comprehensive evaluation, we applied the same process to another pool of features following an omics swap, thereby validating the robustness of the candidate gene selection. Detailed information on the NPRF algorithm can be found in our previous study [16]. Omics swap We swapped the roles of the training gene set and the validation gene set, then applying the feature selection algorithm on the new training set (i.e., the original training gene set excluding the transcriptomics dataset and including the genomics dataset) to calculate the \(\:{S}_{j}^{DR}\) scores for these new features, thereby identifying a new set of candidate biomarkers. Integrative systems biology analysis ( Data systems layer ) To better understand the identified biomarkers, we employed systems biology, a system-level and holistic biological framework, to predict the underlying mechanisms and properties of these biomarkers. The integrative systems biology framework was applied to uncover key DTgenes from the identified biomarkers. The entire process included pathway enrichment analysis, genetic network analysis, and topological analysis. Pathway enrichment analysis To uncover the relevant pathways associated with the identified biomarkers, we employed GO functional annotation ( http://geneontology.org/ ). We extracted soybean pathways and corresponding genes from the GO database, identifying 4,896 pathways that included 48,606 unique genes mapped to the Williams 82 reference genome version 2 ( Glycine max Wm82.a2.v1). GO annotation pathways systematically offer a standardized catalog to classify functional genes based on their biological functions and molecular mechanisms. We established a filtering criterion (≤ 4 genes, or ≥ 1,500 genes) to eliminate pathways with insufficient information or overly common pathways. After filtering, 3,194 pathways including 31,383 unique genes remained. We applied a self-competitive method (hypergeometric test) to discern the significant pathways associated with the biomarkers. The hypergeometric test is a prevalent statistical approach in biology, utilized to assess whether annotated pathways are significantly enriched ( p -value < 0.05) in specific gene sets [40]. The p -value is computed as follows: \(\:p-\text{v}\text{a}\text{l}\text{u}\text{e}=\:\sum\:_{x=k}^{K}\frac{\left(\genfrac{}{}{0pt}{}{K}{x}\right)\left(\genfrac{}{}{0pt}{}{N-K}{m-x}\right)}{\left(\genfrac{}{}{0pt}{}{N}{m}\right)}\) , where N is the number of genes in the whole soybean genome, K is the number of genes in a particular pathway, m is the number of biomarkers, and x is the number of biomarkers in a specific pathway. To mitigate potential FPs, we applied the Bonferroni correction (Bonferroni-corrected p -value ≤ 0.005). Gene network and topological analysis Molecular interactions between biological components are often represented as gene networks, based on graph theory. Graph theory is an academic discipline that demonstrates the relationship between pairwise objects. In our study, we used undirected network to identify functionally related genes. Nodes represented the biomarkers, and edges represented the molecular interaction between the nodes. We utilized functional modules in SoyNet [41] ( https://www.inetbio.org/soynet/ ) to calculate edges by summing log-likelihood scores. Cytoscape [42] v3.10.0 was then used to create the gene network of biomarkers, and we merged two gene networks (merged with swap case). Biomarkers with degree numbers between the 40th and the 80th percentiles were selected for further key DTgenes selection. Key DTgenes discovery To refine the number of DTgenes for precision breeding, we implemented the same systems biology analysis process on biomarkers obtained by omics swapping. Overlapping biomarkers between the two gene networks were selected as key DTgenes. Three selection criteria were then applied to further refine key DTgenes based on their importance index ( \(\:{II}_{j}\) ) and genetic score ( \(\:{S}_{ij}^{G}\) ). DTgenes were removed if they did not meet the following criteria: (1) \(\:{II}_{j}\) is greater than 1.5. (2) \(\:{S}_{ij}^{G}\) is greater than 2 within genomic and transcriptomic data platforms. (3) \(\:{S}_{ij}^{G}\) is greater than 1.5 within proteomic and metabolomic data platforms. The remained DTgenes were regarded as key DTgenes in this study. Key genes validation (evaluation layer) Identifying and validating key DTgenes is essential for developing drought-resistant soybean varieties. By integrating computational and molecular approaches, we can ensure robust and reliable identification of genes that contribute to drought tolerance. The combination of statistical validation, previous research evidence, and molecular experiments ensures a robust identification process for precision breeding of drought-tolerant soybean varieties. Computational methods (1) Statistical validation To evaluate the contribution of candidate genes to phenotypic traits, reliable statistical approaches were employed. We compared key DTgenes with the other DTgenes, intermediate genes, and remaining genes across three validation databases [12, 38, 39]. These databases encompass seven different drought treatments: 6h, 12h, 24h, 5d, 7d, 3–4 weeks, and 6–7 weeks. The Wilcoxon rank-sum test was applied with a significance threshold of p -value < 0.05. (2) Previous research evidence Five biological levels were used to verify key DTgenes in this study: RNA, protein, function, homologs, and gene editing for homologs level. At the RNA level, key DTgenes may exhibit significant expression changes under water-deprivation conditions. At the protein level, we assumed that the corresponding encoded proteins may show changes in abundance under water-restricted environments. At the function level, key DTgenes were expected to be significantly enriched in metabolic pathways relevant to drought tolerance response, as per the PMN database. At the homologs level, we aimed to identify changes in gene or protein patterns of corresponding homologous genes and sought direct evidence for functional annotation via gene editing techniques. We focused on legume crops and model plant Arabidopsis. We utilized SoyBase ( https://www.soybase.org/ ) to search for Arabidopsis homologs with the highest sequence similarity, and the Basic Local Alignment Search Tool (BLAST) to find top hits of well-studied or validated legume crops homologs. Default parameters were used for BLAST searches, with 80–100% as the filtering criteria for query cover and percent identity. Molecular experiments – drought stress treatment Soybean seeds were surface-sterilized in 1% sodium hypochlorite and rinsed with distilled water. Seeds culture and stress conditions followed previous studies with some modifications [43, 44]. Approximately 10–15 soybean seeds were sown in sandy soil in plastic pots (240 mm length × 240 mm width × 190 mm depth). Pots were kept well-watered until the plants had five unfolded trifoliate leaves. Drought stress was imposed by withdrawing water for 1 day and 5 days, with soil moisture calculated by gravimetric water content. Untreated plants were used as controls. Shoots and roots were collected and immediately frozen in liquid nitrogen for RNA extraction. Total RNA was isolated from shoots and roots using TRIzol reagent following the manufacturer’s protocol. Subsequently, 5 µg of total RNA was mixed with 500 ng of oligo(dT) 18 and 200 U Superscript™ III reverse transcriptase (Invitrogen), and the mixture was reverse transcribed at 50°C for 60 min. Real-time PCR was performed using the Luna® Universal qPCR Master Mix (NEB) with the gene-specific primers according to the manufacturer’s protocol. Relative gene expression ratios were calculated using the 2 −ΔΔCT method, with untreated plants cDNA as the reference sample and actin as the reference gene. All experiments were conducted in three biological replicates. Statistical analysis to validate significant differences in key DTgenes transcript quantities under drought stress was performed using one-way ANOVA followed by Tukey’s HSD test method facilitated by IBM SPSS statistics software. p -values less than 0.05 were considered statistically significant. Results The development of a systematic data-driven pipeline for feature engineering (DPFE) to identify key drought-tolerant genes (DTgenes) is a critical step for the advancement of precision breeding (Fig. 1 ). This comprehensive pipeline consists of several layers designed to extract and prioritize important genetic information from diverse data sources, including the data source layer, data processing layer, data integration layer, data prioritization layer, data systems layer, and evaluation layer. The pipeline performs thorough data filtering, quality control, and preservation of biologically valuable data across layers, culminating in the identification of key DTgenes from a large-scale feature pool (Fig. 3 ). Data acquisition and curation A total of 788,696 raw data entries were mined from cloud-based data sources, characterized by their diverse data types and formats (Fig. 4 ; Supplementary data 1). These data comprised: (1) Structured data: 781,492 entries stored exclusively in spreadsheet format. (2) Semi-structured data: 3,034 entries in various formats, including PDFs with extractable tables (1,289 entries), CSV (1,695 entries), and HTML. (3) Quasi-structured data: 55 entries in PDFs with unextractable tables. (4) Unstructured data: 3,665 entries in image format. Additionally, we extracted 4,896 biological pathways along with 48,606 unique soybean genes from the GO website ( http://geneontology.org/ ). Data ETL and Data standardization Data ETL and standardization processes were applied to ensure the consistency and usability of diverse datasets obtained from high-throughput technologies (omics data) and traditional molecular methods (non-omics data). The omics datasets included genomics, transcriptomics, proteomics, and metabolomics, while the non-omics data comprised genomic, transcriptomic, proteomic, metabolomic, and epigenomic data categories. Consequently, five OnO data platforms were established: genomic, transcriptomic, proteomic, metabolomic, and epigenomic data platforms. From structured data formats, a total of 139,428 genomic data points, 577,788 transcriptomic data points, 60,112 proteomic data points, and 4,614 metabolomic data points were extracted. Semi-structured data contributed an additional 1,037 genomic data points, 1,748 transcriptomic data points, 199 proteomic data points, and 50 metabolomic data points. Quasi-structured data yielded 55 genomic data points, while unstructured data provided 3,647 transcriptomic data points and 9 metabolomic and epigenomic data points. The detailed hierarchical information of data types, formats, and OnO data platforms is illustrated in Fig. 4 . Data integration and systematic omics swap A total of 147,590 data points were processed and integrated across five OnO data platforms, encompassing 49,836 genomic, 46,713 transcriptomic, 47,931 proteomic, 3,080 metabolomic, and 30 epigenomic data points. To quantify the association with drought-tolerant responses, we adopted genetic scores ( \(\:{S}_{ij}^{G}\) ) across these platforms and calculated the important index for individual features ( \(\:{II}_{j}\) ). The genetic scores before rescaling exhibited diverse ranges: genomic (0.000007 to 9.78), transcriptomic (0.000000038 to 30.78), proteomic (1 to 16.46), metabolomic (1.33 to 8.9), and epigenomic (1.41 to 2.2). The score distribution for each OnO data platform showed right skewness and was subsequently rescaled into the range of [0, 6], except for epigenomic data platform. Summing up the rescaled \(\:{S}_{ij}^{G}\) across the five OnO data platforms resulted in 55,465 unique features in the training gene set. The ranges of \(\:{II}_{j}\) and \(\:{S}_{j}^{DR}\) were [0.7, 2.0] and [0.000011, 29.11], respectively (Fig. 5 a). To ensure the robustness and comprehensiveness of our pipeline, we conducted a systematic omics swap involving one genomics database and two transcriptomics databases. Following the same procedure, we obtained 106,593 data points, including 3,098 genomic and 52,454 transcriptomic data points. The number of data points remained unchanged in the other OnO data platforms. This process resulted in 54,086 unique features in the training gene set, with \(\:{II}_{j}\) ranging from 0.7 to 1.8 and \(\:{S}_{j}^{DR}\:\) ranging from 0.00000044 to 26.15 (Fig. 5 b). Data prioritization The objective of feature prioritization is to retain the most critical biomarkers for subsequent analyses. To achieve this, we applied a non-parametric random forest (NPRF) algorithm, utilizing gene frequency computation to select features. This process effectively reduced the initial pool of over 55,000 features to a refined set of 135 features (Fig. 6 a). These features, considered as biomarkers, demonstrated significant differences in gene expression levels within the validation gene set (p < 0.00001). Applying the same procedure to the omics swap scenario resulted in the selection of 133 features (Fig. 6 b), which were also identified as biomarkers. These biomarkers showed statistical significance within the validation gene set ( p < 0.05) (Supplementary data 2). Integrative systems biology analysis Pathways enrichment analysis In this study, we first applied a hypergeometric test for pathway enrichment to reveal critical pathways involved in drought-tolerant responses. We identified 21 (omics unswapped scenario) and 24 (omics swap scenario) GO pathways that were significantly associated with drought-tolerant responses based on 135 and 133 biomarkers, respectively. A total of 112 biomarkers were identified in enriched pathways in both scenarios, with 60 overlapping biomarkers, highlighting the complexity of the underlying biological mechanisms. Gene network analysis To elucidate the complex nature of drought tolerance, we utilized the built-in tools within SoyNet to explore multiple pathways and gene interactions between selected biomarkers. We set up two strategies to identify functional modules. First, we constructed two distinct functional gene networks using 135 biomarkers (from omics unswapped scenario) and 133 biomarkers (from omics swap scenario) identified in feature prioritization, comprised 162 nodes and 1,778 edges, and 139 nodes and 1,268 edges, respectively. Second, we selected 112 biomarkers identified in enriched pathways from each scenario for gene network construction, forming complex gene networks comprising 136 nodes and 1,391 edges (from omics unswapped scenario), and 139 nodes and 1,393 edges (from omics swap scenario), respectively. To construct an integrative gene network, we utilized Cytoscape [42] to pool two gene networks constructed by distinct strategies for each scenario, resulting in comprehensive functional interactions comprising 162 nodes and 1,717 edges in the omics unswapped scenario (Fig. 7 a) and 160 nodes and 1,393 edges in the omics swap scenario (Fig. 7 b). Topology analysis We conducted a topology analysis to identify genes with significant degrees of interaction to reveal their potential role in drought tolerance. Degrees of functional interactions for each of two scenarios ranged from 0 to 105 (omics unswapped) and 0 to 97 (omics swap), respectively. We further identified 52 (out of 162 genes) and 65 (out of 160 genes) biomarkers within the 40th to 80th percentile degree numbers between 3–53 (omics unswapped) and 2–45 (omics swap), respectively. Among them, 26 biomarkers overlapped. After eliminating genes with zero-degree numbers, we identified 47 biomarkers (35 biomarkers from the omics unswapped scenario and 21 biomarkers from the omics swap scenario, with 9 overlapped biomarkers) and 53 intermediate genes (33 genes from the omics unswapped scenario and 31 genes from the omics swap scenario, with 11 overlapped genes) that lay outside this range. Key DTgenes discovery We identified 26 biomarkers from the intersection of two gene networks (omics unswapped network and omics swap network) within the 40th and 80th percentile of degree numbers as DTgenes candidates. We applied three selection criteria based on the distribution of \(\:{S}_{ij}^{G}\) and \(\:{II}_{j}\) (please refer to the Materials and Methods section for details) for further refine these candidates. Consequently, 10 candidate DTgenes were identified and regarded as key DTgenes in this study (Fig. 8 ). Evaluations Computational methods (1) Statistical validation The 10 key DTgenes in the corresponding gene networks statistically outperformed (p-values < 0.05) the other biomarkers within gene networks (i.e., other DTgenes), intermediate genes, and the remaining genes across all drought treatment spans (Fig. 9 ). (2) Previous research evidence Multiple lines of evidence were identified across various biological levels (Fig. 10 ). Only one key DTgene, Glyma.05g056300 , has supporting evidence at the RNA level. At the protein level, four key DTgenes ( Glyma.09g178100 , Glyma.11g111400 , Glyma.12g037400 , and Glyma.19g106800 ) were associated with protein abundance research in response to drought treatment. Their encoded protein expression has been proven to be upregulated during drought treatment. At the functional level, four key DTgenes ( Glyma.06g015900 , Glyma.11g111400 , Glyma.12g037400 , and Glyma.19g106800 ) were identified within the Plant Metabolic Network (PMN) database ( https://plantcyc.org/ ), and were associated with five shared metabolic pathways, including Glycolysis I (from glucose-6-phosphate), Glycolysis II (from fructose-6-phosphate), Glycolysis IV (plant cytosol), Calvin-Benson-Bassham cycle, and sucrose biosynthesis I (from photosynthesis). These metabolic pathways are related to glycolysis, photosynthesis, and carbohydrates metabolism. Six key DTgenes demonstrated evidence at the homologs level ( Glyma.05g056300 , Glyma.06g015900 , Glyma.07g043600 , Glyma.11g111400 , Glyma.12g037400 , and Glyma.19g106800 ). Among them, Glyma.07g043600 and Glyma.19g106800 lacked evidence in legume crops and Arabidopsis, respectively. All top hits evidence of legume crops were found in Pisum sativum (pea) with query cover and percent identity over 99% and 86%, respectively. Five key DTgenes were found to have evidence in Arabidopsis through gene editing for homologs level ( Glyma.05g056300 , Glyma.06g015900 , Glyma.11g111400 , Glyma.12g037400 , and Glyma.16g211700 ). These were proved to affect plant growth under drought conditions through gene knockout techniques in Arabidopsis ( Glyma.05g056300 ) or to lose crucial physiological mechanisms associated with drought tolerance response ( Glyma.06g015900 , Glyma.11g111400 , Glyma.12g037400 , and Glyma.16g211700 ). No evidence was observed in soybean and model plants for two key DTgenes, Glyma.09g163900 and Glyma.13g341500 . Molecular experiments We performed qRT-PCR to examine the expression of 10 key DTgenes under drought-induced stress to verify the findings of the transcriptome data of DTgenes (Fig. 11 ). In the current study, the expression of five genes ( Glyma.06g015900 , Glyma.07g043600 , Glyma.09g163900 , Glyma.11g111400 , and Glyma.12g037400 ) in soybean shoot and root was considerably upregulated after withdrawing water for 5 days. The Glyma.05g056300 in the shoot was upregulated at 1 day and downregulated at 5 days, while in the root was significantly upregulated from 1 to 5 days. The highest expression of Glyma.09g178100 in the shoot was found at 5 days and in the root at 1 day. A pattern of downregulation was observed for Glyma.13g341500 . One gene ( Glyma.16g211700 ) was induced at 1 day, then slightly down at 5 days. Finally, no significant difference was observed in the shoot for Glyma.19g106800 , but the highest level was found at 5 days in the root system. Discussion Drought presents a significant challenge in agriculture, consistently leading to the most substantial losses among various abiotic stresses in crop yields and economic impacts globally. In 2019, drought events resulted in a 7% reduction in soybean production, translating to over 12 billion US dollars in global losses [2, 45]. Given the crucial role of soybean in food, agriculture, industry, and health sectors, these figures highlight the urgent need to develop strategies to mitigate the impacts of drought. Conventional methods for identifying important biomarkers are time-consuming and labor/resource-intensive [46–48] and often unable to keep pace with the rapid environmental changes induced by climate change. Moreover, the multifaceted nature of drought-tolerant responses, involving complex interactions between various genetic and phenotypic traits [49], requires a more holistic approach. In this study, we proposed that integrating diverse biological data using a comprehensive DPFE can enhance the discovery of key DTgenes in soybean. This method ensures that only biologically valuable data are preserved and analyzed, increasing the likelihood of identifying robust DTgenes. Despite potential challenges such as data heterogeneity and computational complexity, the comprehensive design of pipeline and rigorous quality control on data enhance its effectiveness. By integrating structured, semi-structured, quasi-structured, and unstructured data, we ensured comprehensive coverage of relevant information, while advanced feature extraction techniques enhance the reliability of the findings. The subsequent data standardization and mapping techniques facilitated the alignment of heterogeneous data sources, crucial for meaningful analysis. The development of DPFE aims to not only reduce the cost and timeline for plant breeding but also enhance precision and applicability of key genes identification. By using resources from previous research (multi-omics and non-omics datasets), data-driven techniques, and systems biology, our pipeline streamlines the process of key DTgenes selection, ensuring an effective approach to enhance drought tolerance in soybean crops. This approach addresses current limitations in single-study analyses and sets a precedent for future research in crop improvement under changing climatic conditions. We conducted extensive data mining and curation to deal with data heterogeneity (distinct data formats and types) and data quality (presence of false positives (FPs) and noise, data incompleteness, and irrelevant data), ending up with 788,696 genetic raw data. Managing large amounts of data remains challenging due to the presence of FPs and noise. To address these problems, we developed tailored strategies for information retrieval from each data type and conducted data quality assurance at each layer to ensure data accuracy and reliability. By acquiring and preprocessing datasets specifically related to drought responses and organizing them based on their structural characteristics, we streamlined the subsequent data processing and analysis steps, leading to more accurate and meaningful insights into soybean drought tolerance mechanisms. The data ETL process is crucial in deriving meaningful biological information from multiple data types and formats. The variation in data formats introduced complexity into data ETL process, increasing the difficulty and time required for data standardization. This process ensures that data from diverse sources can be reliably compared and analyzed, leading to more accurate and meaningful biological insights. Raw data extracted from quasi-structured PDFs and images required more time and manual efforts in data processing to convert and standardize the data to uniform formats [50]. Integration of non-structured and structured data is a significant task in the ETL process to ensure data integrity. Continuous monitoring, validation, and refinement of the ETL process are crucial. The integration of data from multiple data platforms and multidisciplinary research was crucial to minimize the underlying risks of FPs, bias, and noise. Unlike traditional approaches that may rely on single-omics or solitary molecular technique, our approach incorporates feature selection with a conceptual integration strategy across five OnO data platforms. Numerous studies have demonstrated that a comprehensive model containing multiple OnO data can compensate for uncertain or missing data in any unbalanced dataset, outperforming models including only single-omics alone [51, 52]. This holistic approach is based on technological advancements and large amounts of previous research, revolutionizing our understanding of soybean genetics for agronomic traits. Although non-omics data were more inconsistent and variable, posing challenges for data integration [52], a well-designed conceptual integration strategy can cope with underlying FPs and compensate for the noise and bias of each data platform. The conceptual integration strategy allowed a more flexible structure to determine the mathematical approaches for fitting the existing dataset’s patterns and research objectives compared to the statistical and model-based integration strategy [37], enhancing the potential for advancing our understanding of drought-tolerant responses in soybean. To effectively integrate and analyze this multifaceted data, we developed a scoring scheme, \(\:{II}_{j}\) , and normalization process. By rescaling the score of data points within different data platforms into a unified numerical score [0, 6], we ensured a common scale across entire OnO data platforms, addressing the inherent variability in molecular properties. In terms of the \(\:{II}_{j}\) index, which increases the sensitivity for biomarker identification, by increasing (or decreasing) the weight of features with strong or consistent (or weak or inconsistent) evidence, the system ensures that only the most credible candidates are highlighted as key DTgenes in the subsequent analysis. Through this strategy, we can improve the high FP rate of biomarkers findings in traditional breeding progress. The systematic omics swap involved a genomics database and two transcriptomics databases. The primary reason for selecting the genomics database as the swapped database was its highly comprehensive genome data, containing over 130,000 SNPs mapped to over 44,000 soybean reference genes. This approach produced an alternative gene pool, enhancing the comprehensiveness of features for the identifying biomarkers more likely to confer drought tolerance in soybean. The application of the NPRF algorithm for feature selection was pivotal in the current study. We strategically determined the number of features by the distribution of gene frequency, setting the selection threshold just before the first significant drop in frequency. The feature selection process reduced the feature pool from over 55,000 to 135 features, ensuring the meaningful and manageable biomarker candidates. The combination of ETL process, conceptual integration, and feature selection facilitated fast and precise biomarker discovery. The integrative approach combining multiple omics datasets and systematic omics swap provides a powerful framework for robust biomarker discovery. The consistency observed in the omics swap scenario further validates the reliability of our feature selection method. The rationale behind the omics swap procedure is to mitigate the risk of overfitting and ensure that the identified biomarkers are not specific to a particular subset of the data. This method allows for cross-validation of key genes and reduces the risk of FPs, enhancing the reliability of our findings. The insights gained from this integrative analysis can significantly contribute to advancing precision breeding strategies aimed at improving drought tolerance in soybeans. The surge of biological data has fundamentally shifted the paradigm of systems biology in plants [53]. Through integrative pathway enrichment and construction of gene networks, we have successfully obtained novel insights into the molecular interactions and functional implications of complex regulatory networks governing drought-tolerant responses in soybean. Gene network and topological analysis emerged as complex yet informative, highlighting critical nodes and connections key to the adaptation to drought in soybean. Biomarkers involved with common biological processes are likely to be functionally related. Genes with higher connectivity (degree number), however, do not necessarily imply greater functional relevance. Genes with intermediate connectivity are considered to have the highest functional co-annotation rate [54, 55]. To find the most functionally related biomarkers, we eliminated biomarkers with a degree number ranking below the 40th percentile, as they contribute minimally to the network. Conversely, we also excluded biomarkers with a degree number above the 80th percentile to avoid overly general biomarkers with lower specific functional relationships. This method allowed us to focus on biomarkers most likely integral to drought tolerance in soybean. By combining two gene networks, 10 key DTgenes were identified. These key DTgenes potentially represent the pivotal role in the soybean drought-tolerant responses, supported by high functional relationship within gene networks and enriched pathways. However, no key DTgenes were identified in the epigenomic data platform due to limited relevant information. Nine out of ten key DTgenes ( Glyma.05g056300 , Glyma.06g015900 , Glyma.07g043600, Glyma.09g163900 , Glyma.09g178100 , Glyma.11g111400 , Glyma.12g037400, Glyma.13g341500 , Glyma.16g211700 , Glyma.19g106800 ) were recorded across four OnO data platforms, except for Glyma.07g043600 , which was recorded in three data platforms (genomic, transcriptomic, and proteomic). Of these key DTgenes, the highest frequency was identified in Glyma.06g015900 , indicating its involvement in protein-protein interactions and biological processes. Despite Glyma.07g043600 being recorded in only three platforms, it showed significant scores in genomic, transcriptomic, and proteomic data, suggesting its importance in drought stress response. The integration of large-scale multiple OnO data has provided a valuable resource for identifying potential key DTgenes. Nevertheless, the reliability and functionality of these genes must be rigorously assessed through further research. We expect that these genes will exhibit significant changes across various biological levels. Legume crops were selected as model plant species due to their close phylogenetic relationship, facilitating the inference of functions for key DTgenes. Not all key DTgenes, however, have identifiable homologous genes or functions in legume crops. To bridge this gap, Arabidopsis was introduced as a complementary model species due to its well-annotated functional genome. Sequence identity in BLAST is the widely-known strategy for identification of homologous proteins and genes [56]. We searched for corresponding Arabidopsis homologs in SoyBase based on the highest sequence similarity associated with key DTgenes. In this study, we successfully identified ten key DTgenes in soybean using a data-driven pipeline, with eight having corresponding evidence in previous research and two ( Glyma.09g163900 and Glyma.13g341500 ) lacking relevant supporting documents. Among the identified key DTgenes, Glyma.05g056300 emerged as a prime candidate, showing strong evidence across multiple levels, including RNA, homologs, and gene editing in Arabidopsis. This gene encodes acid-soluble invertase in soybean (GmA-INV / GmA-VINV1.1A) and exhibits significant expression changes during drought at key flowering stages, suggesting a pivotal role in drought response [57]. Its orthologs in pea ( Pisum sativum ) [58] ( PsVINV1.1 ) and Arabidopsis [59] ( At1g12240 / AtVIN2 ) have shown similar regulation, indicating drought resistance. Stomatal conductance is significantly reduced in VIN2 -suppressed Arabidopsis plants, which are less tolerant to water deficit than wild-type plants, reinforcing its importance [60]. Glyma.11g111400 and Glyma.12g037400 , encoding fructose biphosphate aldolase 2 (GmFBA2), also demonstrated substantial evidence across four biological levels. There genes are homologous to an identical Arabidopsis gene ( At4g38970 ) and are involved in critical metabolic pathways, especially under drought stress conditions in soybean cotyledons [61]. The Arabidopsis homolog ( AtFBA2 ) plays a central role in photosynthesis and energy metabolism [62] and is essential for starch synthesis in the Calvin cycle [63]. The fba2 mutant shows growth retardation due to inhibited sugar synthesis [63]. Modifications in carbon metabolism demonstrate drought improvement in resistant cultivars, storing more carbon for energy generation to maintain physiological mechanisms under drought [64]. Additionally, the Arabidopsis homologs AtFBA2 ( KIW84_074187 and KIW84_074188 ) were observed in peas. Their encoded proteins are significantly upregulated under drought conditions, playing crucial roles in glycolytic process in pea seedlings [65]. Glyma.06g015900 , encoding glyceraldehyde-3-phosphate dehydrogenase B subunit (GAPB), is homologous to At1g42970 ( AtGAPB ). AtGAPB shows significant expression changes under different drought durations [66]. The gapb mutant transgenic line demonstrates a significant decrease in CO 2 assimilation in photosynthesis, implying GAPB’s potential importance in regulating drought tolerance [67, 68]. Orthologous genes of Glyma.06g015900 were identified in pea, showing significant changes under water stress and osmotic stress recovery stages [69]. Other key DTgenes such as Glyma.09g178100 and Glyma.19g106800 , showed strong evidence in soybean [61, 70], while no other evidence was found in homologs for Glyma.09g178100 . The orthologs of Glyma.19g106800 in pea encodes glyceraldehyde-3-phosphate dehydrogenase A and shows significant protein abundance changes under water stress [69]. In previous transcriptomic research, Glyma.19g106800 was one of the top 50 most highly expressed transcripts under drought stress [71]. Both Glyma.07g043600 and Glyma.16g211700 were supported by homologous genes in Arabidopsis [70, 72]. The homolog of Glyma.07g043600 in Arabidopsis ( At1g54050 ) is indirectly associated with drought tolerance, with significant gene expression changes under oxidative stress. Oxidative stress is induced by multiple abiotic and biotic stresses, including drought [72]. Glyma.16g211700 is homologous to At1g73260 (Kunitz-type trypsin inhibitor 1, AtKTI1 ) in Arabidopsis, which causes early flowering when overexpressed [73], a behavior associated with drought escape [74]. Previous proteomics studies indicate that proteins related to energy and photosynthesis play crucial roles in drought tolerance [70], aligning with our functional level discoveries. Four key DTgenes were found to be involved in glycolysis, carbohydrates metabolism, and the Calvin cycle. The remaining key DTgenes, Glyma.09g163900 and Glyma.13g341500 , represent novel discoveries. While they lacked prior research evidence, their significant expression level differences across three OnO data platforms highlight their potential as new contributors to drought tolerance in soybean. These two genes are innovative discoveries made possible by our comprehensive, multi-faceted analytical approach. Further analysis using ShinyGO v0.80 [75] identified six genes ( Glyma.05g056300 , Glyma.06g015900 , Glyma.09g178100 , Glyma.11g111400 , Glyma.12g037400 , and Glyma.19g106800 ) enriched in eight GO terms related to carbohydrate metabolism. Two genes ( Glyma.05g056300 and Glyma.09g178100 ) were significantly enriched in five KEGG pathways, mainly related to carbohydrate metabolism and secondary metabolite synthesis. Four KEGG pathway networks demonstrated the roles of Glyma.05g056300 and Glyma.09g178100 in these networks (Fig. 12 ). These findings suggest that improving or optimizing carbohydrate metabolism in soybean may confer drought tolerance to soybean seedlings. The DPFE framework used in the current study, which combines data ETL and feature selection, not only identified key genes consistent with previous research but also uncovered novel genes that may play crucial roles in drought-tolerant responses. To further confirm the functionality of these genes, future work should investigate their association with drought tolerance through RNA silencing or gene knockout. Environmental stress, including drought, significantly restricts plant growth and development. Increasing crop tolerance to drought is an effective strategy for maintaining yields [76]. In this study, most key DTgenes showed increased expression, while one gene exhibited decreased expression following drought stress, according to qRT-PCR data. Numerous plant mechanisms for environmental stress, including our earlier research [44], have demonstrated such up-and-down regulation. Three genes ( Glyma.05g056300 , Glyma.11g111400 , and Glyma.12g037400 ) related to carbohydrate transport and metabolism [77–80] were upregulated, indicating that soybeans, like other plants, respond to drought stress through carbohydrate metabolism. Among the 10 key DTgenes, Glyma.07g043600 exhibited the highest relative gene expression. Predicted as GmHSP20, this gene was highly induced in RNA-seq data [81]. Plant heat shock proteins (Hsps) typically regulate antioxidant enzymes positively, improving membrane integrity and detoxifying reactive oxygen species during stress tolerance [82]. Glyma.09g163900 and Glyma.13g341500 , which lacked previous evidence, showed significantly different expression patterns under drought treatment in this study, suggesting they may be novel biomarkers for drought tolerance research in soybean. Thus, our qRT-PCR data may indicate the potential biological roles of key DTgenes in plant protection against drought stress. Despite the promising potential of our data-driven pipeline, several challenges in the present study must be acknowledged. Firstly, our findings relied entirely on existing research, and the completeness and quality of these data could introduce noise and biases. To address this, we implemented a data ETL process with stringent standards for data quality control and management to minimize these risks. Integrating multiple data platforms based on different molecular methodologies can minimize the publication bias and FPs, reducing the risk of biased results. Secondly, the heterogeneity of data sources poses significant integration difficulties. Our data-driven pipeline employs several strategies, including data ETL and standardization, multi-platform integration, data mapping and alignment, advanced computational tools, and systematic omics swap, to address these issues effectively. Thirdly, the availability and accessibility of relevant datasets can be challenging, as some datasets may be dispersed across different repositories and require considerable effort to locate and retrieve. Our pipeline incorporates strategies for data aggregation and retrieval, data integration and standardization, and centralized repository setup to manage this challenge effectively. Finally, the applicability of this pipeline may be limited during the data integration stages due to the scarcity of experimental data in a domain. As the domain develops and more data accumulates, the discovery of biomarkers through this pipeline will become more consistent and stable, further supporting its reliability, robustness, and precision. Additionally, processing and analyzing large-scale datasets require robust infrastructure and sophisticated algorithms. The data-driven pipeline for feature engineering presented in this study represents a significant advancement in the field of crop genetics and breeding. By leveraging the power of big data and systems biology, we identified key DTgenes that hold promise for improving drought tolerance in soybean. This approach addresses the limitations of traditional breeding methods and sets the stage for future research and applications in various agricultural contexts. Declarations 1. Ethics approval and consent to participate This article does not contain any studies with human participants or animals performed by any of the authors. 2. Consent for publication Not applicable. This article does not contain any studies with human participants or animals performed by any of the authors that require consent for publication. 3. Availability of data and materials The raw data acquired from source of article underlying this article are available in Github ( github.com/edwardkao6413/Raw-data-of-soybean-drought-tolerant-research.git ), further inquiries or any data support can be directed to the corresponding author. 4. Competing interests The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. 5. Funding This study is funded by Chung Cheng Agriculture Science and Social Welfare Foundation, FUTURETECH ELECTRONICS CO., LTD, National Science and Technology Council (NSTC 112-2622-B-005-001), and Advanced Plant and Food Crop Biotechnology Center from The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education in Taiwan. 6. Authors' contributions C.F.Kao conceived the study conception and design. P.H.Kao, C.W.Li, C.W.Tseng, and C.F.Kao collected the data. P.H.Kao and S.Baiya performed the data analysis. P.H.Kao, S.Baiya, C.F.Kao, S.Y.Chen, and Y.H.Huang interpreted the data and results. P.H.Kao, S.Baiya, and C.F.Kao drafted the manuscript. C.F.Kao and P.H.Kao revised the manuscript. All authors read and approved the final manuscript. 7. Acknowledgements This study was granted by Chung Cheng Agriculture Science and Social Welfare Foundation and FUTURETECH ELECTRONICS CO., LTD. We extend our appreciation to the National Science and Technology Council (NSTC 112-2622-B-005-001). For their financial support, which contributed (in part) to the success of this research. Furthermore, we would like to thank the Advanced Plant and Food Crop Biotechnology Center from The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education in Taiwan. Their support has been invaluable in advancing agricultural research and promoting scientific innovation. References Hodges T, French V. Soyphen: Soybean Growth Stages Modeled from Temperature, Daylength, and Water Availability. J Agron. 1985;77:500–5. Razzaq A, Wani SH, Saleem F, Yu M, Zhou M. Rewilding crops for climate resilience: economic analysis and de novo domestication strategies. J Exp Bot. 2021;72:6123–39. Kramer PJ. J. S. Boyer. Water relations of plants and soils. San Diego: Academic; 1995. Seleiman MF, Al-Suhaibani N, Ali N, Akmal M. Alotaibi. Drought stress impacts on plants and different approaches to alleviate its adverse effects. Plants. 2021;10:259. Desclaux D, Roumet P. Impact of drought stress on the phenology of two soybean (Glycine max L. Merr) cultivars. Field Crops Res. 1996;46:61–70. Lafitte H, Yongsheng G, Yan S, Li Z. Whole plant responses, key processes, and adaptation to drought stress: the case of rice. J Exp Bot. 2007;58:169–75. Yadav RS, Hash C, Bidinger F, Devos K. Howarth. Genomic regions associated with grain yield and aspects of post-flowering drought tolerance in pearl millet across stress environments and tester background. Euphytica. 2004;136:265–77. Byrum J, Davis C, Doonan G, Doubler T. D. Foster. Advanced analytics for agricultural product development. Interfaces. 2016;46:5–17. Kusmec A, Zheng Z, Archontoulis S, Ganapathysubramanian B, Hu G. Interdisciplinary strategies to enable data-driven plant breeding in a changing climate. One Earth. 2021;4:372–83. Ahmar S, Gill RA, Jung K-H, Faheem A, Qasim MU. Conventional and molecular techniques from simple breeding to speed breeding in crop plants: recent advances and future outlook. Int J Mol Sci. 2020;21:2590. Sharma N, Raman H, Wheeler D, Kalenahalli Y, Sharma R. Data-driven approaches to improve water-use efficiency and drought resistance in crop plants. Plant Sci. 2023;336:111852. Saleem A, Roldán-Ruiz I, Aper J, Muylle H. Genetic control of tolerance to drought stress in soybean. BMC Plant Biol. 2022;22:1–19. Weckwerth W, Ghatak A, Bellaire A, Chaturvedi P. Varshney. PANOMICS meets germplasm. Plant Biotechnol J. 2020;18:1507–25. Ward RM, Schmieder R, Highnam G, Mittelman D. Big data challenges and opportunities in high-throughput sequencing. Syst Biomed. 2013;1:29–34. Nekrutenko A, Taylor JJNRG. Next-generation sequencing data interpretation: enhancing reproducibility and accessibility. Nat Rev Genet. 2012;13:667–72. Kao P-H, Baiya S, Lai Z-Y, Huang C-M. Jhan. An advanced systems biology framework of feature engineering for cold tolerance genes discovery from integrated omics and non-omics data in soybean. Front Plant Sci. 2022;13:1019709. Bell G, Hey T. Szalay. Beyond the data deluge. Science. 2009;323:1297–8. Pal S, Mondal S, Das G, Khatua S. Ghosh. Big data in biology: The hope and present-day challenges in it. Gene Rep. 2020;21:100869. Isaac NJ, van Strien AJ, August TA, de Zeeuw MP. Roy. Statistics for citizen science: extracting signals of change from noisy ecological data. Methods ecol evol. 2014;5:1052–60. R. M. Dorazio. Accounting for imperfect detection and survey bias in statistical analysis of presence-only data. Glob Ecol Biogeogr. 2014;23:1472–84. Fithian W, Elith J, Hastie T, Keith DA. Bias correction in species distribution models: pooling survey and collection data for multiple species. Methods ecol evol. 2015;6:424–38. Isaac NJ, Jarzyna MA, Keil P, Dambly LI. Boersch-Supan. Data integration for large-scale models of species distributions. Trends Ecol Evol. 2020;35:56–67. Larose DT, Larose CD. An Introduction to Data Mining of 1. In: editor. An Introduction to Data Mining of Book. New Jersey: Wiley; 2014. Vijayarani S, Ilamathi MJ. Nithya. Preprocessing techniques for text mining-an overview. IJCSC. 2015;5:7–16. Sultana A. Unraveling the Data Structures of Big Data, the HDFS Architecture and Importance of Data Replication in HDFS. Int J Res Eng Technol. 2018;5. Dhaouadi A, Bousselmi K, Gammoudi MM, Monnet S. Hammoudi. Data warehousing process modeling from classical approaches to new trends: Main features and comparisons. Data. 2022;7:113. Zebari R, Abdulazeez A, Zeebaree D, Zebari D. Saeed. A comprehensive review of dimensionality reduction techniques for feature selection and feature extraction. J Appl Sci Technol Trends. 2020;1:56–70. El-Sappagh SHA, Hendawi AMA. El Bastawissy. A proposed model for data warehouse ETL processes. J KING SAUD UNIV-COM. 2011;23:91–104. Sheth BP, Thaker VS. Plant systems biology: insights, advances and challenges. Planta. 2014;240:33–54. Hong W-J, Kim Y-J, Chandran AKN. Jung. Infrastructures of systems biology that facilitate functional genomic study in rice. RICE. 2019;12:1–17. Pazhamala LT, Kudapa H, Weckwerth W, Millar AH. Varshney. Systems biology for crop improvement. Plant Genome. 2021;14:e20098. Fukushima A, Kusano M, Redestig H, Arita M, Saito K. Integrated omics approaches in plant systems biology. Curr Opin Chem Biol. 2009;13:532–8. Fukushima A, Kanaya S, Nishida K. Integrated network analysis and effective tools in plant systems biology. Front Plant Sci. 2014;5:598. Yoosefzadeh Najafabadi M, Hesami M, Eskandari M. Machine learning-assisted approaches in modernized plant breeding programs. Genes. 2023;14:777. Shin JH, Vaughn JN, Abdel-Haleem H, Chavarro C. Abernathy. Transcriptomic changes due to water deficit define a general soybean response and accession-specific pathways for drought avoidance. BMC Plant Biol. 2015;15:1–13. Tamang BG, Li S, Rajasundaram D, Lamichhane S, Fukao T. Overlapping and stress-specific transcriptomic and hormonal responses to flooding and drought in soybean. Plant J. 2021;107:100–17. Hahne F, Huber W, Gentleman R, Falcon S. Falcon. Hypergeometric testing used for gene set enrichment analysis of 14. In: editor. Hypergeometric testing used for gene set enrichment analysis of Book. New York: Springer; 2008. pp. 207–20. Kim E, Hwang S, Lee I. SoyNet: a database of co-functional networks for soybean Glycine max. Nucleic Acids Res. 2017;45:D1082–9. Saito R, Smoot ME, Ono K, Ruscheinski J. Wang. A travel guide to Cytoscape plugins. Nat Methods. 2012;9:1069–76. Mutava RN, Prince SJK, Syed NH, Song L, Valliyodan B. Understanding abiotic stress tolerance mechanisms in soybean: A comparative evaluation of soybean response to drought and flooding stress. Plant Physiol Biochem. 2015;86:109–20. Jhan L-H, Yang C-Y, Huang C-M, Lai M-C. Huang. Integrative pathway and network analysis provide insights on flooding-tolerance genes in soybean. Sci Rep. 2023;13:1980. Kim W, Iizumi T, Nishimori M. Global patterns of crop production losses associated with droughts from 1983 to 2009. J Appl Meteorol Climatol. 2019;58:1233–44. Du W, Wang M, Fu S, Yu D. Mapping QTLs for seed yield and drought susceptibility index in soybean (Glycine max L.) across different environments. J GENET GENOMICS. 2009;36:721–31. Du W, Yu D, Fu S. Detection of quantitative trait loci for yield and drought tolerance traits in soybean using a recombinant inbred line population. J Integr Plant Biol. 2009;51:868–78. Specht J, Chase K, Macrander M, Graef G, Chung J. Soybean response to water: a QTL analysis of drought tolerance. Crop Sci. 2001;41:493–509. Shelake RM, Kadam US, Kumar R, Pramanik D, Singh AK. Engineering drought and salinity tolerance traits in crops through CRISPR-mediated genome editing: Targets, tools, challenges, and perspectives. Plant Commun. 2022;3. Zhang W, Zeng Y, Jiao M, Ye C, Li Y. Integration of high-throughput omics technologies in medicinal plant research: The new era of natural drug discovery. Front Plant Sci. 2023;14:1073848. López de Maturana E, Alonso L, Alarcón P, Martín-Antoniano IA. Pineda. Challenges in the integration of omics and non-omics data. Genes. 2019;10:238. Greener JG, Kandathil SM, Moffat L, Jones DT. A guide to machine learning for biologists. Nat Rev Mol Cell Biol. 2022;23:40–55. Hou J, van Leeuwen J, Andrews BJ, Boone C. Genetic network complexity shapes background-dependent phenotypic expression. Trends Genet. 2018;34:578–86. Siegal ML, Promislow DE, Bergman A. Functional and evolutionary inference in gene networks: does topology matter? Genetica. 2007;129:83–103. W. R. Pearson. An introduction to sequence similarity (homology) searching. Curr. Protoc. Bioinform. 2013;42:3.1. 1-3.1. 8. Du Y, Zhao Q, Chen L, Yao X, Zhang H. Effect of drought stress during soybean R2–R6 growth stages on sucrose metabolism in leaf and seed. Int J Mol Sci. 2020;21:618. Morin A, Kadi F, Porcheron B, Vriet C, Maurousset L. Genome-wide identification of invertases in Fabaceae, focusing on transcriptional regulation of Pisum sativum invertases in seed subjected to drought. Physiol Plant. 2022;174:e13673. Ni DA. Role of vacuolar invertase in regulating Arabidopsis stomatal opening. Acta Physiol Plant. 2012;34:2449–52. Chen S-F, Liang K, Yin D-M, Ni D-A. Zhang. Ectopic expression of a tobacco vacuolar invertase inhibitor in guard cells confers drought tolerance in Arabidopsis. J Enzyme Inhib Med. 2016;31:1381–5. Wang X, Oh M, Sakata K, Komatsu S. Gel-free/label-free proteomic analysis of root tip of soybean over time under flooding and drought stresses. J Proteom. 2016;130:42–55. Zadražnik T, Hollung K, Egge-Jacobsen W, Meglič V. Šuštar-Vozlič. Differential proteomic analysis of drought stress response in leaves of common bean (Phaseolus vulgaris L). J Proteom. 2013;78:254–72. Lu W, Tang X, Huo Y, Xu R, Qi S. Identification and characterization of fructose 1, 6-bisphosphate aldolase genes in Arabidopsis reveal a gene family with diverse responses to abiotic stresses. Gene. 2012;503:65–74. Cuellar-Ortiz SM, De La Paz M, Arrieta‐Montiel J, Acosta‐Gallegos AA. Covarrubias. Relationship between carbohydrate partitioning and drought resistance in common bean. PLANT CELL ENVIRON. 2008;31:1399–409. Julia S, Daria G, Veronika P, Alexander T, Alena K. Protein glycation and drought response of pea (Pisum sativum L.) root nodule proteome: A proteomics approach. Commun Biol. 2021;66:210–24. Zeng L, Deng R, Guo Z, Yang S, Deng X. Genome-wide identification and characterization of Glyceraldehyde-3-phosphate dehydrogenase genes family in wheat (Triticum aestivum). BMC Genom. 2016;17:1–10. Feller U. Drought stress and carbon assimilation in a warming climate: Reversible and irreversible impacts. J Plant Physiol. 2016;203:84–94. Alqurashi M. Investigating the Role of RAP2. 12 and GAPDH in Photosynthetic CO2 Fixation to Improve Drought Tolerance [Doctoral dissertation]. Place: Institution; 2019. Brosowska-Arendt W, Gallardo K, Sommerer N, Weidner S. Changes in the proteome of pea (Pisum sativum L.) seeds germinating under optimal and osmotic stress conditions and subjected to post-stress recovery. Acta Physiol Plant. 2014;36:795–807. Yu X, James A, Yang A, Jones A, Mendoza-Porras O. A comparative proteomic study of drought-tolerant and drought-sensitive soybean seedlings under drought stress. Crop Pasture Sci. 2016;67:528–40. Chen W, Yao Q, Patil GB, Agarwal G, Deshmukh RK. Identification and comparative analysis of differential gene expression in soybean leaf tissue under drought and flooding stress revealed by RNA-Seq. Front. Plant Sci. 2016;7:1044. Dat J, Vandenabeele S, Vranova E, Van Montagu M, Inzé* D. Dual action of the active oxygen species during plant stress responses. Cell Mol Life Sci. 2000;57:779–95. Kim J-g, Baek S-A, Im K-H. Overexpression of a Kunitz-type trypsin inhibitor (AtKTI1) causes early flowering in Arabidopsis. Plant Growth Regul. 2009;59:75–81. Boerma HR. J. E. Specht. Soybeans: improvement, production and uses. American Society of Agronomy; 2004. Ge SX, Jung D, Yao R. ShinyGO: a graphical gene-set enrichment tool for animals and plants. Bioinformatics. 2020;36:2628–9. Shan L, Xu Y, Wu D, Hu J. T. Yu. Effects of Salicylic Acid on Growth, Physiology, and Gene Expression in Rice Seedlings Under Salt and Drought Stress. Plant Stress. 2024;100413. Yu X, Fu X, Yang Q, Jin H, Zhu L. Genome-wide variation analysis of four vegetable soybean cultivars based on re-sequencing. Plants. 2021;11:28. Kumar V, Goyal V, Mandlik R, Kumawat S, Sudhakaran S. Pinpointing genomic regions and candidate genes associated with seed oil and protein content in soybean through an integrative transcriptomic and QTL meta-analysis. Cells. 2022;12:97. Su L, Xu C, Zeng S, Su L, Joshi T. Large-Scale Integrative Analysis of Soybean Transcriptome Using an Unsupervised Autoencoder Model. Front. Plant Sci. 2022;13:831204. Schläpfer P, Zhang P, Wang C, Kim T, Banf M. Genome-wide prediction of metabolic enzymes, pathways, and gene clusters in plants. Plant Physiol. 2017;173:2041–59. Xiao R, Han Q, Liu Y, Zhang X, Hao Q. Melatonin Attenuates the Urea-Induced Yields Improvement Through Remodeling Transcriptome and Rhizosphere Microbial Community Structure in Soybean. Front microbiol. 2022;13:903467. Lone RA, Sarvendra K, Singh V, Bano N, Bag SK. Adaptation of winged bean (Psophocarpus tetragonolobus (L.) DC.) to drought stress is mediated by root-tuber heat-shock proteins and specific metabolites. Curr Plant Biol. 2022;32:100266. Additional Declarations No competing interests reported. Supplementary Files Supplementarydata.zip Cite Share Download PDF Status: Published Journal Publication published 19 Mar, 2025 Read the published version in Journal of Big Data → Version 1 posted Editorial decision: Revision requested 20 Dec, 2024 Reviews received at journal 20 Dec, 2024 Reviews received at journal 20 Dec, 2024 Reviewers agreed at journal 18 Dec, 2024 Reviewers agreed at journal 18 Dec, 2024 Reviews received at journal 12 Dec, 2024 Reviews received at journal 12 Dec, 2024 Reviewers agreed at journal 04 Dec, 2024 Reviewers agreed at journal 04 Dec, 2024 Reviewers invited by journal 02 Sep, 2024 Editor assigned by journal 16 Aug, 2024 Submission checks completed at journal 26 Jul, 2024 First submitted to journal 23 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4791767","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":343391094,"identity":"5c405dd2-07c6-46dd-8ea1-acb79afc11fc","order_by":0,"name":"Pei-Hsiu Kao","email":"","orcid":"","institution":"Melbourne Integrative Genomics, The University of Melbourne","correspondingAuthor":false,"prefix":"","firstName":"Pei-Hsiu","middleName":"","lastName":"Kao","suffix":""},{"id":343391095,"identity":"52dd1cf5-7dc9-4f57-9e4f-c6ce3c350aa1","order_by":1,"name":"Supaporn Baiya","email":"","orcid":"","institution":"Department of Resource and Environment Faculty of Science at Sriracha Kasetsart University","correspondingAuthor":false,"prefix":"","firstName":"Supaporn","middleName":"","lastName":"Baiya","suffix":""},{"id":343391096,"identity":"7e7333e4-d4fd-435b-bdb1-8af84f18fdeb","order_by":2,"name":"Chong-Wei Lee","email":"","orcid":"","institution":"Department of Agronomy, National Chung Hsing University","correspondingAuthor":false,"prefix":"","firstName":"Chong-Wei","middleName":"","lastName":"Lee","suffix":""},{"id":343391097,"identity":"62194a26-79e0-4fd3-96c2-33f96a394c3d","order_by":3,"name":"Chia-Wen Tseng","email":"","orcid":"","institution":"Department of Agronomy, National Chung Hsing University","correspondingAuthor":false,"prefix":"","firstName":"Chia-Wen","middleName":"","lastName":"Tseng","suffix":""},{"id":343391098,"identity":"23957509-dd62-4928-9692-a0dd6de9296b","order_by":4,"name":"Shu-Yun Chen","email":"","orcid":"","institution":"Department of Agronomy, National Chung Hsing University","correspondingAuthor":false,"prefix":"","firstName":"Shu-Yun","middleName":"","lastName":"Chen","suffix":""},{"id":343391099,"identity":"d69e4d38-510d-4fc3-818a-c2762bffd3b5","order_by":5,"name":"Yen-Hsiang Huang","email":"","orcid":"","institution":"Department of Agronomy, National Chung Hsing University","correspondingAuthor":false,"prefix":"","firstName":"Yen-Hsiang","middleName":"","lastName":"Huang","suffix":""},{"id":343391100,"identity":"d3d90bac-340c-4c9b-8d50-9a93d7aa6e01","order_by":6,"name":"Chung-Feng Kao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuUlEQVRIiWNgGAWjYFACxocPPjaAGY0HiNTCbGw4s4FBAqilgWgtZsK8YC0MDMRpkXc/zMZsu8OmTrf9MNCWGptogloMzySzPc49kyZhdiYRqOVYWm4DQS0z+I8b57YdljA7ANTC2HCYGC3MbNKWIC3nHxKpRV4CqIURpOUGsbYY8CQzG/a2pUluuwG0JYEYv8i3H2Z88LPNht/sfPrDBx9qbIiw5QAyL4GQcrAtBA0dBaNgFIyCUQAAA/dDnD5jzSUAAAAASUVORK5CYII=","orcid":"","institution":"Department of Agronomy, National Chung Hsing University","correspondingAuthor":true,"prefix":"","firstName":"Chung-Feng","middleName":"","lastName":"Kao","suffix":""}],"badges":[],"createdAt":"2024-07-24 02:23:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4791767/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4791767/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s40537-025-01078-w","type":"published","date":"2025-03-19T15:57:59+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":63013221,"identity":"4ffb20ae-3958-46b6-8cb0-9230d3bb8a7f","added_by":"auto","created_at":"2024-08-22 06:11:03","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":524331,"visible":true,"origin":"","legend":"\u003cp\u003ePipeline for the discovery of key drought-tolerant genes (DTgenes).\u003c/p\u003e","description":"","filename":"image1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/66897fb63f10a7aab7d370df.jpeg"},{"id":63013871,"identity":"71746e93-6a47-4bde-a9e0-afe1ae99f2e0","added_by":"auto","created_at":"2024-08-22 06:19:03","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":476497,"visible":true,"origin":"","legend":"\u003cp\u003eConceptual framework for noise, bias, and false positive data filtering in the current data-driven pipeline.\u003c/p\u003e","description":"","filename":"image2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/4dcec65dcaf5edd0f266c8b9.jpeg"},{"id":63013869,"identity":"6f83273b-6062-4302-b710-c14a16eeaab1","added_by":"auto","created_at":"2024-08-22 06:19:03","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":345638,"visible":true,"origin":"","legend":"\u003cp\u003eThe sequential process of narrowing down candidate genes across the pipeline.\u003c/p\u003e","description":"","filename":"image3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/3762db2bc17bbaa872fa21f0.jpeg"},{"id":63013230,"identity":"ddc32376-bd25-484a-9bec-8d253040b0f3","added_by":"auto","created_at":"2024-08-22 06:11:04","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":306393,"visible":true,"origin":"","legend":"\u003cp\u003eHierarchical organization of processed data. Various data formats were mined from the cloud, and OnO data were derived from the diversity of these formats. Each number indicates the quantity of data points processed.\u003c/p\u003e","description":"","filename":"image4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/c16bf845314640ed37328ba4.jpeg"},{"id":63014519,"identity":"1d74800c-9d6f-421c-bb60-e3c0a2d7c2c0","added_by":"auto","created_at":"2024-08-22 06:27:03","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":203817,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of drought-tolerant response scores (𝑆\u003csub\u003e\u003cem\u003ej\u003c/em\u003e\u003c/sub\u003e\u003csup\u003e\u003cem\u003eDR\u003c/em\u003e\u003c/sup\u003e) for (a) the omics unswapped scenario (55,465 OnO features) and (b) the omics swap scenario (54,086 OnO features). Each black dot represents one biomarker.\u003c/p\u003e","description":"","filename":"image5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/22e974c31cef0bfa4ca7f2b5.jpeg"},{"id":63013224,"identity":"20fb3c15-f5e3-402b-a83d-a7493e20a353","added_by":"auto","created_at":"2024-08-22 06:11:03","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":105385,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of gene frequency of selected biomarkers for (a) the omics unswapped scenario and (b) the omics swap scenario.\u003c/p\u003e","description":"","filename":"image6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/ed124458ddc058112a3afcf8.jpeg"},{"id":63013874,"identity":"bc36ea4e-b8ad-4422-b4bf-43d3e8f88394","added_by":"auto","created_at":"2024-08-22 06:19:04","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":582072,"visible":true,"origin":"","legend":"\u003cp\u003eGene networks of biomarkers for (a) the omics unswapped scenario and (b) the omics swap scenario.\u003c/p\u003e","description":"","filename":"image7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/f7768c3bbf5ce226fa570966.jpeg"},{"id":63013227,"identity":"ec7eec3e-29bc-4977-a21a-35c2c6f5b03b","added_by":"auto","created_at":"2024-08-22 06:11:04","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":450721,"visible":true,"origin":"","legend":"\u003cp\u003ePhysical positions of biomarkers, overlapping biomarkers, and key DTgenes on soybean chromosomes for (a) the omics unswapped scenario and (b) the omics swap scenario.\u003c/p\u003e","description":"","filename":"image8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/a29762590c9cadd3acbb4ddb.jpeg"},{"id":63014522,"identity":"79a03063-ec43-43ac-a8c3-b5cfec360322","added_by":"auto","created_at":"2024-08-22 06:27:06","extension":"jpeg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":234402,"visible":true,"origin":"","legend":"\u003cp\u003eValidation study of the key DTgenes compared to the other biomarkers, the intermediate genes, and the remaining genes using an independent RNA-seq data. (* \u003cem\u003ep\u003c/em\u003e \u0026lt; 0.05; ** \u003cem\u003ep\u003c/em\u003e\u0026lt; 0.01; *** \u003cem\u003ep\u003c/em\u003e \u0026lt; 0.001)\u003c/p\u003e","description":"","filename":"image9.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/3bed427144f10e39fdb09ed4.jpeg"},{"id":63013229,"identity":"ed2bab78-f16e-4a8f-af99-91c39daa3b97","added_by":"auto","created_at":"2024-08-22 06:11:04","extension":"jpeg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":381987,"visible":true,"origin":"","legend":"\u003cp\u003eDecagonal representation of supporting evidence for each key DTgene across multiple biological levels.\u003c/p\u003e","description":"","filename":"image10.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/8e354c0826cb90919d25092f.jpeg"},{"id":63013875,"identity":"57218c97-0acf-4d04-b11e-6a8be1fa9073","added_by":"auto","created_at":"2024-08-22 06:19:04","extension":"jpg","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":105294,"visible":true,"origin":"","legend":"\u003cp\u003eGene expression results for molecular experiments under 1- and 5-day drought treatments. The blue bars represent shoot samples, and the orange bars represent root samples.\u003c/p\u003e","description":"","filename":"image11.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/7e690c0639b257afee4ee1b3.jpg"},{"id":63013233,"identity":"18fefaf4-afc6-4dfe-9311-ec86163fd200","added_by":"auto","created_at":"2024-08-22 06:11:04","extension":"jpeg","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":543553,"visible":true,"origin":"","legend":"\u003cp\u003eThe roles of two genes enriched in KEGG pathways within the pathway networks: (a) Cyanoamino acid metabolism, (b) Biosynthesis of various plant secondary metabolites, (c) Galactose metabolism, and (d) Starch and sucrose metabolism. Red marks represent the key DTgenes, with 3.2.1.21 and 3.2.1.26 denoting \u003cem\u003eGlyma.09g178100\u003c/em\u003e and \u003cem\u003eGlyma.05g056300\u003c/em\u003e, respectively.\u003c/p\u003e","description":"","filename":"image12.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/a6aa73b76378fbabc9f548fd.jpeg"},{"id":79120503,"identity":"0a0a7ed1-ecfe-48f0-8e08-48c1955f1490","added_by":"auto","created_at":"2025-03-24 16:09:03","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5161338,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/becb38d1-6f87-46db-a1d4-3c340fca5763.pdf"},{"id":63013234,"identity":"da88fa67-7a9b-42ad-a113-5cca09be1444","added_by":"auto","created_at":"2024-08-22 06:11:06","extension":"zip","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":45112617,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarydata.zip","url":"https://assets-eu.researchsquare.com/files/rs-4791767/v1/e23dccbbe26e7006868f87c8.zip"}],"financialInterests":"No competing interests reported.","formattedTitle":"Unveiling insights into drought-tolerant responses in soybean: a data-driven pipeline for feature engineering in biomarker discovery","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSoybean (\u003cem\u003eGlycine max\u003c/em\u003e L.) is an important global food crop owing to its high protein and oil content, but it faces high susceptibility to environmental fluctuations, particularly water availability [1]. Extreme weather caused by the climate change incurs the escalating of drought events, which poses a substantial threat to soybean cultivation and further reduce agricultural productivity [2, 3]. Drought stress primarily results from water scarcity and inadequate soil moisture supply, inducing various adverse effects on crop physiology [4]. Persistent water deficit during the vegetative growth forces early transition to reproductive stage, leading to substantial yield losses [5]. Furthermore, drought stress during early reproductive or grain-filling stages incurs pollen sterility and impaired grain set, exacerbating economic losses [6, 7].\u003c/p\u003e \u003cp\u003eConventional breeding is time-consuming practice for developing new crop varieties [8]. The time constraints, inherent genetic stochasticity, environmental fluctuations, and substantial resource investments set formidable obstacles to breeding program [9, 10]. The rapid pace of climate change frequently surpasses the timeline for developing new and resilient varieties, exacerbating agricultural losses. To overcome these issues, there is a growing demand for more rapid, accurate, and cost-efficient methodologies in identifying key resistant genes for tolerant varieties development. A data-driven approach emerges as a paradigm-shifting solution by applying extensive biological datasets. The integration of multidisciplinary research data and systems biology approaches provides novel insights into crop improvements by the use of comprehensive information, facilitating the breeding progress [11].\u003c/p\u003e \u003cp\u003eThe current understanding of drought tolerance in soybean still remains limited. Although some advanced high-throughput techniques have been developed and applied to biomarkers discovery for drought improvement, it still remains challenges due to the complex interplay of multiple biological processes in drought-tolerant response. Drought-tolerant response cannot be explored and interpreted relying solely on a single study or unilateral molecular approach, such as genome-wide association study (GWAS) or RNA-seq, which may not capture enough information. For instance, current GWAS research can only account for approximately 30% of the phenotypic variance observed between drought-tolerant lines and commercial lines. This suggests that the contribution of each single nucleotide polymorphism (SNP) to the variance explainability is exceedingly low [12, 13]. Other non-omics or non-whole genome studies also face challenges for biomarkers discovery. The variations in materials, methods, and environmental conditions applied in individual study incurs issues in standardization and reproducibility of the experiment outcome [14, 15]. To expand our understanding to the remaining phenotypic variance, there is an urgent need for integrative and interdisciplinary approaches. Integrative approaches, such as combining omics and non-omics (OnO) data, have shown promise in uncovering the genetic mechanisms underlying complex traits in various organisms [16], indicating their potential applicability in soybean drought tolerance research.\u003c/p\u003e \u003cp\u003eThe raising of high-throughput techniques, propelled by advancements in computational capabilities and software tools, creates a new era for biology [17]. This paradigm shift poses challenges in identification of important information from the soaring volume of biological data. The integration of multi-dimensional biological data presents an opportunity to gain comprehensive insights into the molecular mechanisms underlying complex traits in soybean. This integration process, however, encounters certain challenges, primarily due to the heterogeneity of data types [18], requiring specialized tools for effective handling. Structured omics data offers standardized formats and well-defined variables, facilitating analysis and interpretation. On the other hand, unstructured data sources, including textual information from scientific literature and phenotypic data from field trials, lack a predefined structure, posing challenges in data organization and integration. Furthermore, variability in data quality and completeness across different datasets can introduce biases and noise [19]. Despite these challenges, the integration of multi-dimensional biological data holds immense promise for advancing soybean breeding and genomics research by providing a holistic understanding of the genetic factors governing agronomic traits [20\u0026ndash;22].\u003c/p\u003e \u003cp\u003eData mining plays a critical role in the extraction of meaningful information from large-scale datasets [23]. A large challenge pertains to the abundance of non-structured data, which constitute over 80% of valuable data pool [24, 25]. Text mining stands out as particularly significant for wanted information extracted from diverse data types, including structured, semi-structured, quasi-structured, and unstructured formats. Structured data, characterized by predefined fields and formats, including relational databases and spreadsheets, whereas unstructured data, such as images, text files, audio files, and PDFs, lacks a rigid organizational schema. Semi-structured data, represented by formats like comma-separated value (CSV), extensible markup language (XML), and hypertext markup language (HTML), occupies an intermediate position between structured and unstructured data, with part of organizational properties but lacking uniformity in database structure [26]. Quasi-structured data, although less frequently discussed, the irregular formats requires specialized tools for efficient extraction and processing [27]. A comprehensive understanding of diverse data types is indispensable in bioinformatics, facilitating the effective utilization of the wide spectrum of data essential for our research endeavors.\u003c/p\u003e \u003cp\u003eData extraction-transformation-loading (ETL) serves as a fundamental process for the cleansing, customization, and reformatting of data within the realm of data warehousing [28]. By using data-driven approaches, the ETL process enables the systematic extraction of relevant information from diverse biological databases, including OnO data. Initially, raw data is extracted from various sources databases, along with textual information from scientific literature and field trials. Subsequently, data cleaning techniques are applied to remove inconsistencies, errors, and irrelevant data [29, 30], ensuring the quality and integrity of the dataset [30]. This systematic approach is crucial for maintaining data quality, particularly considering the non-standardized nature of biological data [31]. Data filtering procedures are then implemented to focus on specific features relevant to drought tolerance traits in soybean. Additionally, data mapping techniques facilitate the integration and alignment of heterogeneous data sources, allowing for the correlation of genetic information with phenotypic traits. Following this, data sorting algorithms are employed to organize the dataset according to predefined criteria in data repository or warehouse, facilitating efficient analysis and interpretation. Finally, the processed data is loaded into computational tools and databases for further analysis and feature selection.\u003c/p\u003e \u003cp\u003eSystems biology, including pathway enrichment analysis, network analysis, and topology properties, represents a powerful tool to understand cellular and organismal functions holistically at a systemic level [32]. Through the integration of diverse OnO data and the application of mathematical models, systems biology offers a comprehensive approach to analyzing and interpreting heterogeneous datasets, thereby enhancing our understanding of plant physiology and biological systems [33, 34]. By aggregating information from multiple genes into annotated pathways, such as Gene Ontology (GO) terms, the joint effects of numerous genes can be assessed to identify enriched biological pathways [35] that are modulated in response to a specific trait. Network analysis allows for the construction of interaction networks among genes, proteins, and other molecular entities involved in the trait, revealing the intricate interactions between genes and pathways [36]. Furthermore, topology-based analysis, such as degree, centrality measures, connectivity, and clustering coefficients, can effectively model biological phenomena by incorporating pathways to identify key genes within the network that play crucial roles in the trait. Together, these approaches collectively offer a systems-level perspective of the molecular landscape associated with a specific trait, facilitating the prioritization of key genes and pathways for further investigation and targeted breeding strategies.\u003c/p\u003e \u003cp\u003eIn the current study, we introduced a novel data-driven pipeline for key genes discovery, which aimed to overcome the limitations of conventional breeding practices for crop improvement. This innovative strategy harnesses the flexibility of data-driven methodologies, combining diverse data from multi-disciplinary research. It provides a more comprehensive and efficient approach to address the complex genotype-phenotype associations to support decision-making in breeding programs. The proposed data-driven pipeline not only offers a promising solution for identifying key genes for crop improvement but also serves as a protocol for future research and applications across various fields.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eA data-driven pipeline for feature engineering\u003c/h2\u003e \u003cp\u003eWe devised a data-driven pipeline for feature engineering, termed DPFE, and aimed to identify key genes associated with soybean drought-tolerant responses. The DPFE system comprises six layers: a data source layer involving cloud-based data mining and wrangling, a data processing layer encompassing data extraction, transformation, loading, and harmonization to retrieve wanted information from diverse data types and enhance data quality, a data integration layer where processed data is integrated into training and validation gene sets, a data prioritization layer utilizing feature selection to identify biomarkers, a data systems layer employing pathway enrichment, gene networks, and topological analysis for key DTgenes discovery, and an evaluation layer assessing key DTgenes through computational methods and molecular experiments. The comprehensive architecture of DPFE is depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eData mining and wrangling\u003c/b\u003e \u003cb\u003e(\u003c/b\u003e\u003cb\u003edata source layer\u003c/b\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo comprehensively elucidate the genetic mechanisms underlying drought-tolerant responses in soybean, we initiated an extensive data acquisition process across multiple sources, including data warehouses, data lakes, and cloud-based experimental repositories. Relevant datasets were systematically retrieved using targeted search terms on platforms such as NCBI, focusing specifically on drought tolerance-related data. This included genomics, transcriptomics, proteomics, metabolomics, and epigenomics data pertinent to soybean under drought conditions. Following data mining, rigorous data wrangling protocols were implemented to ensure data quality and relevance. The protocols involved data curation and categorization. We implemented strict curation protocols to filter and retain only data directly relevant to drought-tolerant responses. This involved retaining datasets from field trials with controlled soil water content and experiments simulating drought conditions using varying levels of polyethylene glycol. These protocols were crucial to not only align the data to our research objective by laying the foundation onto the data of biological relevance and integrity, but also minimize the noise and maximize the signal-to-noise ratio. To effectively manage the diversity of data, we categorized datasets based on their structural and formatting characteristics. This included distinguishing between structured data (e.g., spreadsheets) and non-structured data. Non-structured data incorporated semi-structured data (e.g., HTML, CSVs, PDFs with extractable tables), quasi-structured data (e.g., PDFs without extractable tables), and unstructured data (e.g., images). This categorization was pivotal for organizing the training gene set and guiding subsequent data processing and analysis methods tailored to each data type.\u003c/p\u003e \u003cp\u003e \u003cb\u003eData ETL process\u003c/b\u003e \u003cb\u003e(\u003c/b\u003e\u003cb\u003edata processing layer\u003c/b\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003cp\u003eWe devised a comprehensive data-driven ETL process combined with robust feature engineering techniques to identify key genes associated with soybean drought-tolerant responses. The ETL process is a powerful procedure to retrieve wanted information from multiple data sources, maximizing the utility of available data for more accurate and insightful discoveries. This process is vital for organizing and preparing data, ensuring its quality and relevance to inform our analysis.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eExtraction\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe employed Python v3.11.4 for data extraction across various sources including structured, semi-structured, quasi-structured, and unstructured data. Structured data was processed with Pandas 2.0.3 for spreadsheets. Semi-structured data was accessed through BeautifulSoup v4.21.1, Pandas 2.0.3, and Tabula v1.0.5 for web scraping from HTML, CSV files, and PDF files with extractable tables, respectively. Quasi-structured data extraction required Tabula v1.0.5 and extensive manual curation for non-extractable PDF tables. Unstructured data, particularly images, underwent manual annotation to identify relevant features. Semi-structured and quasi-structured data required customized scripts for processing, whereas unstructured data involved more intensive manual curation due to its complexity.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eTransformation\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFollowing extraction, data underwent a rigorous transformation process. This included cleaning, filtering, and standardizing to eliminate irrelevant or erroneous data. Structured data was reformatted and categorized based on its relevance to soybean drought-tolerant response. Non-structured data was converted into structured formats, enabling seamless integration into our data platforms. We meticulously categorized the transformed data into appropriate OnO data platforms to maintain clarity.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eLoading\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eData loading involved systematic integration of the transformed and cleaned data into a structured data repository. This repository serves as a centralized storage system, allowing for efficient data retrieval and analysis. Ensuring data integrity and consistency at this stage is critical for subsequent analysis. To maintain consistency, we introduced standardized columns to align the new data with the existing data structure. The standardized columns include data type (structured, semi-structured, quasi-structured, unstructured), data format (spreadsheet, HTML, CSV, PDF, image), OnO (omics, non-omics), system component level (genome, transcriptome, proteome, metabolome, epigenome), feature (SNP, gene, SSR, QTL), chromosome, start genomic position, end genomic position, genome version, drought duration, statistical significance value (p-value), measure of linkage strength (logarithm of the odds, LOD), measure of differential expression (FC), computed value representing feature importance (score), developmental stage of the plant (growth stage), and source or reference of the data (reference). These standardized columns ensure that all data entries adhere to a uniform format, facilitating easier data manipulation and analysis. A final quality assurance step is conducted to ensure that the data in the repository meets the required standards for completeness, accuracy, and relevance. Any issues identified during this step are addressed before the data is made available for analysis.\u003c/p\u003e \u003cp\u003e \u003cb\u003eMultiple feature integration\u003c/b\u003e \u003cb\u003e(\u003c/b\u003e\u003cb\u003eData integration layer\u003c/b\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe integration of multiple data features is fundamental to capture the complexity of biological systems. By combining different data features, we can achieve a more comprehensive understanding of the genotype-phenotype associations that underpin drought-tolerant responses in soybean. In this study, we addressed the challenge of integration of heterogeneous data from various omics (OnO) level of the system components (denoted as data platforms), which include diverse data types and originate from different molecular methodologies, plant materials, and environmental conditions. Recognizing the inherent risk of FPs in molecular biology experiments, our primary objective was to minimize this risk in the discovery of key DTgenes. Our hypothesis was that biomarkers genuinely associated with drought-tolerant responses would exhibit consistent patterns across various OnO data platforms, while irrelevant biomarkers, such as FPs, would not. Those FPs would neither achieve high scores across data platforms nor reach top-ranking positions, making them less likely to be selected as DTgenes candidates (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). To tackle this, we employed conceptual integration along with integrative feature selection strategy based on the principle that the whole outperforms the sum of its parts [37]. This approach involved harmonizing and normalizing different data points across multiple OnO data platforms using a scoring scheme. Here, we partitioned the entire dataset into the training gene set and the validation gene set.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eMultifaceted feature integration (constructing the training gene set)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe developed a comprehensive approach for assessing drought-tolerant responses in soybean through the integration of multiple features. Specifically, we established a drought-tolerant response score (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{j}^{DR}\\)\u003c/span\u003e\u003c/span\u003e) for each feature in the training gene set calculated as the product of the sum of genetic scores across multiple OnO data platforms (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{.j}^{G}\\)\u003c/span\u003e\u003c/span\u003e) and their respective importance index (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e) to represent the magnitude of association between the feature and the trait. The genetic score \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e was determined using a score transformation function, denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{h}^{ij}(\\bullet\\:)\\)\u003c/span\u003e\u003c/span\u003e, applied to synthesize a single or several data points from each OnO data platform (indexed by \u003cem\u003ei\u003c/em\u003e) for a specific feature (indexed by \u003cem\u003ej\u003c/em\u003e). We summed \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e across all OnO data platforms and subsequently rescaled it to a standardized range between 0 and 6. To evaluate the significance of each feature in relation to the trait, we introduced the important index (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e). Initially, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e was assigned a value of 1.0 to represent a neutral starting point for each feature and adjusted based on its performance threshold (\u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05, LOD\u0026thinsp;\u0026gt;\u0026thinsp;1.5, or FC\u0026thinsp;\u0026gt;\u0026thinsp;1.5) within individual OnO data platform. Specifically, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e increased by 0.2 if the feature met the threshold, and decreased by 0.1 if it did not. Features with indeterminate performance maintained their initial \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e value. This approach allows us to systematically quantify the importance of each feature in contributing to soybean drought tolerance. By integrating data from multiple platforms and considering their respective significance, we aim to identify key genes associated with drought-tolerant responses in soybean.\u003c/p\u003e \u003cp\u003eIn our methodology, we utilized diverse scoring schemes and data normalization techniques specific to each OnO data platform to calculate the genetic score \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{.j}^{G}\\)\u003c/span\u003e\u003c/span\u003e for individual features (indexed by \u003cem\u003ej\u003c/em\u003e). Given the heterogeneous nature of the data, including various types of numerical genetic information such as \u003cem\u003ep\u003c/em\u003e-values (for statistical significance in association mapping data), LOD scores (for assessing the strength of genetic linkage in linkage or association mapping data), and FC values (for evaluating the magnitude of expression differences in expression data), derived from molecular characteristics within each OnO data platform, it was essential to employ platform-specific procedures for score calculation. To ensure comparability and consistency across platforms, we implemented appropriate data normalization techniques tailored to the characteristics of each dataset. These normalization procedures aimed to account for inherent variations in data distributions and experimental conditions, thereby facilitating meaningful comparisons and integrations across multiple platforms.\u003c/p\u003e \u003cp\u003eWe employed various transformation functions to combine multiple \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\pi\\:\\)\u003c/span\u003e\u003c/span\u003e values obtained from \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:m\\)\u003c/span\u003e\u003c/span\u003e independent hypothesis tests \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:({H}_{0}:{\\pi\\:}_{s}\\sim\\text{U}\\left[\\text{0,1}\\right],\\:s=1,\\cdots\\:,m)\\)\u003c/span\u003e\u003c/span\u003e into a single test statistic for the \u003cem\u003ej\u003c/em\u003e-th feature within the \u003cem\u003ei\u003c/em\u003e-th OnO data platform. Specifically, we used the following transformation functions:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{h}_{P}^{ij}\\left(\\bullet\\:\\right)=-2\\sum\\:_{s=1}^{m}\\text{ln}({\\pi\\:}_{s})\\)\u003c/span\u003e \u003c/span\u003e for combining \u003cem\u003ep\u003c/em\u003e-values, which is distributed as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{x}_{2m}^{2}\\)\u003c/span\u003e\u003c/span\u003e under the null hypothesis.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{h}_{NP}^{ij}\\left(\\bullet\\:\\right)=\\text{m}\\text{i}\\text{n}\\left\\{{\\pi\\:}_{1},{\\pi\\:}_{2},\\cdots\\:,{\\pi\\:}_{m}\\right\\}\\)\u003c/span\u003e \u003c/span\u003e for combining non-\u003cem\u003ep\u003c/em\u003e-value genetic information, which follows a \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{B}\\text{e}\\text{t}\\text{a}(1,m)\\)\u003c/span\u003e\u003c/span\u003e distribution under the null hypothesis.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eTo process non-\u003cem\u003ep\u003c/em\u003e-value data (such as LOD scores, FC values, degrees, and pathway scores), we employed a reciprocal of the logarithmically transformed absolute value of various genetic information values to convert these values into \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\pi\\:\\)\u003c/span\u003e\u003c/span\u003e values. This transformation is mathematically represented as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{10}^{-\\text{log}\\left(\\right|.\\left|\\right)}\\)\u003c/span\u003e\u003c/span\u003e, which ensures that the values are normalized on a comparable scale. The genetic score \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{.j}^{G}\\)\u003c/span\u003e\u003c/span\u003e was then calculated using the formula \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{-\\text{l}\\text{o}\\text{g}\\{h}^{ij}(\\bullet\\:)\\}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{h}^{ij}\\left(\\bullet\\:\\right)=\\text{m}\\text{i}\\text{n}\\left\\{{h}_{P}^{ij}\\left(\\bullet\\:\\right),{\\:h}_{NP}^{ij}\\left(\\bullet\\:\\right)\\:\\right\\}\\)\u003c/span\u003e\u003c/span\u003e. A higher value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{.j}^{G}\\)\u003c/span\u003e\u003c/span\u003e suggested a stronger potential \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{.j}^{G}\\)\u003c/span\u003e\u003c/span\u003econnection with drought-tolerant responses in soybean. These transformation and normalization procedures were designed to facilitate meaningful comparisons and integrations across multiple platforms, ensuring that the diverse genetic information could be accurately and consistently interpreted. We designated this integrated dataset as the training gene set.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eValidation gene set\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe selected two omics datasets, which were excluded from the training gene set, as the validation gene set. The first selected transcriptomics dataset contains 104,022 gene expression profiles, encompassing three different drought treatments (6h, 12h, and 24h) conducted on R2 soybean seedlings [38]. The second selected transcriptomics dataset contains 112,094 gene expression profiles, encompassing three different drought treatments (5 days and 7 days) conducted on V1 soybean seedlings [39].\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eOmics swap\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe \u0026lsquo;omics swap\u0026rsquo; procedure is a validation technique designed to ensure the robustness of selected biomarkers by testing their consistency across different data subsets. This involves alternating the roles of a training gene set and a validation set within the integrated feature analysis framework. In this study, we selected one genomics dataset as the new validation set. This dataset contains 139,428 SNPs, mapped to 49,662 genes, with soybean seedlings subjected to 3\u0026ndash;4 weeks and 6\u0026ndash;7 weeks of drought treatment [12]. The goal is to verify that the identified features maintain their significance and predictive power irrespective of the data subset used.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFeature selection and prioritization\u003c/b\u003e \u003cb\u003e(\u003c/b\u003e\u003cb\u003edata prioritization layer\u003c/b\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eFeature prioritization\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eFeature selection and prioritization is a crucial step in identifying key biomarkers associated with drought tolerance responses in soybean. To achieve this, we employed the NPRF algorithm on the training gene set, which is capable of handling high-dimensional data and proficient at selecting a subset of features strongly linked to drought-tolerant responses. The NPRF algorithm constructs decision trees using bootstrap resampling and random permutation. This approach enables the calculation of the frequency of each biomarker by comparing the average scores of bootstrap samples with permutation samples through the Wilcoxon rank-sum test. Specifically, we conducted 1,000 bootstrap iterations and performed 10 rounds of random permutation within each bootstrap iteration to compute the gene frequencies. An appropriate threshold (\u003cem\u003ek\u003c/em\u003e) was then determined, corresponding to a dramatical drop in gene frequency, to identify the most significant features. The top \u003cem\u003ek\u003c/em\u003e features with the highest drought-tolerant response score (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{j}^{DR})\\)\u003c/span\u003e\u003c/span\u003e were selected as our initial set of candidate biomarkers. To ensure a comprehensive evaluation, we applied the same process to another pool of features following an omics swap, thereby validating the robustness of the candidate gene selection. Detailed information on the NPRF algorithm can be found in our previous study [16].\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eOmics swap\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe swapped the roles of the training gene set and the validation gene set, then applying the feature selection algorithm on the new training set (i.e., the original training gene set excluding the transcriptomics dataset and including the genomics dataset) to calculate the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{j}^{DR}\\)\u003c/span\u003e\u003c/span\u003e scores for these new features, thereby identifying a new set of candidate biomarkers.\u003c/p\u003e \u003cp\u003e \u003cb\u003eIntegrative systems biology analysis\u003c/b\u003e \u003cb\u003e(\u003c/b\u003e\u003cb\u003eData systems layer\u003c/b\u003e\u003cb\u003e)\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo better understand the identified biomarkers, we employed systems biology, a system-level and holistic biological framework, to predict the underlying mechanisms and properties of these biomarkers. The integrative systems biology framework was applied to uncover key DTgenes from the identified biomarkers. The entire process included pathway enrichment analysis, genetic network analysis, and topological analysis.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003ePathway enrichment analysis\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eTo uncover the relevant pathways associated with the identified biomarkers, we employed GO functional annotation (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://geneontology.org/\u003c/span\u003e\u003cspan address=\"http://geneontology.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). We extracted soybean pathways and corresponding genes from the GO database, identifying 4,896 pathways that included 48,606 unique genes mapped to the Williams 82 reference genome version 2 (\u003cem\u003eGlycine max\u003c/em\u003e Wm82.a2.v1). GO annotation pathways systematically offer a standardized catalog to classify functional genes based on their biological functions and molecular mechanisms. We established a filtering criterion (\u0026le;\u0026thinsp;4 genes, or \u0026ge;\u0026thinsp;1,500 genes) to eliminate pathways with insufficient information or overly common pathways. After filtering, 3,194 pathways including 31,383 unique genes remained.\u003c/p\u003e \u003cp\u003eWe applied a self-competitive method (hypergeometric test) to discern the significant pathways associated with the biomarkers. The hypergeometric test is a prevalent statistical approach in biology, utilized to assess whether annotated pathways are significantly enriched (\u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05) in specific gene sets [40]. The \u003cem\u003ep\u003c/em\u003e-value is computed as follows: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:p-\\text{v}\\text{a}\\text{l}\\text{u}\\text{e}=\\:\\sum\\:_{x=k}^{K}\\frac{\\left(\\genfrac{}{}{0pt}{}{K}{x}\\right)\\left(\\genfrac{}{}{0pt}{}{N-K}{m-x}\\right)}{\\left(\\genfrac{}{}{0pt}{}{N}{m}\\right)}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cem\u003eN\u003c/em\u003e is the number of genes in the whole soybean genome, \u003cem\u003eK\u003c/em\u003e is the number of genes in a particular pathway, \u003cem\u003em\u003c/em\u003e is the number of biomarkers, and \u003cem\u003ex\u003c/em\u003e is the number of biomarkers in a specific pathway. To mitigate potential FPs, we applied the Bonferroni correction (Bonferroni-corrected \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026le;\u0026thinsp;0.005).\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eGene network and topological analysis\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eMolecular interactions between biological components are often represented as gene networks, based on graph theory. Graph theory is an academic discipline that demonstrates the relationship between pairwise objects. In our study, we used undirected network to identify functionally related genes. Nodes represented the biomarkers, and edges represented the molecular interaction between the nodes. We utilized functional modules in SoyNet [41] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.inetbio.org/soynet/\u003c/span\u003e\u003cspan address=\"https://www.inetbio.org/soynet/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to calculate edges by summing log-likelihood scores. Cytoscape [42] v3.10.0 was then used to create the gene network of biomarkers, and we merged two gene networks (merged with swap case). Biomarkers with degree numbers between the 40th and the 80th percentiles were selected for further key DTgenes selection.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eKey DTgenes discovery\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eTo refine the number of DTgenes for precision breeding, we implemented the same systems biology analysis process on biomarkers obtained by omics swapping. Overlapping biomarkers between the two gene networks were selected as key DTgenes. Three selection criteria were then applied to further refine key DTgenes based on their importance index (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e) and genetic score (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e). DTgenes were removed if they did not meet the following criteria: (1) \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e is greater than 1.5. (2) \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e is greater than 2 within genomic and transcriptomic data platforms. (3) \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e is greater than 1.5 within proteomic and metabolomic data platforms. The remained DTgenes were regarded as key DTgenes in this study.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eKey genes validation (evaluation layer)\u003c/h2\u003e \u003cp\u003eIdentifying and validating key DTgenes is essential for developing drought-resistant soybean varieties. By integrating computational and molecular approaches, we can ensure robust and reliable identification of genes that contribute to drought tolerance. The combination of statistical validation, previous research evidence, and molecular experiments ensures a robust identification process for precision breeding of drought-tolerant soybean varieties.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eComputational methods\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e(1) Statistical validation\u003c/p\u003e \u003cp\u003eTo evaluate the contribution of candidate genes to phenotypic traits, reliable statistical approaches were employed. We compared key DTgenes with the other DTgenes, intermediate genes, and remaining genes across three validation databases [12, 38, 39]. These databases encompass seven different drought treatments: 6h, 12h, 24h, 5d, 7d, 3\u0026ndash;4 weeks, and 6\u0026ndash;7 weeks. The Wilcoxon rank-sum test was applied with a significance threshold of \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003cp\u003e(2) Previous research evidence\u003c/p\u003e \u003cp\u003eFive biological levels were used to verify key DTgenes in this study: RNA, protein, function, homologs, and gene editing for homologs level. At the RNA level, key DTgenes may exhibit significant expression changes under water-deprivation conditions. At the protein level, we assumed that the corresponding encoded proteins may show changes in abundance under water-restricted environments. At the function level, key DTgenes were expected to be significantly enriched in metabolic pathways relevant to drought tolerance response, as per the PMN database. At the homologs level, we aimed to identify changes in gene or protein patterns of corresponding homologous genes and sought direct evidence for functional annotation via gene editing techniques. We focused on legume crops and model plant Arabidopsis. We utilized SoyBase (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.soybase.org/\u003c/span\u003e\u003cspan address=\"https://www.soybase.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to search for Arabidopsis homologs with the highest sequence similarity, and the Basic Local Alignment Search Tool (BLAST) to find top hits of well-studied or validated legume crops homologs. Default parameters were used for BLAST searches, with 80\u0026ndash;100% as the filtering criteria for query cover and percent identity.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eMolecular experiments \u0026ndash; drought stress treatment\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eSoybean seeds were surface-sterilized in 1% sodium hypochlorite and rinsed with distilled water. Seeds culture and stress conditions followed previous studies with some modifications [43, 44]. Approximately 10\u0026ndash;15 soybean seeds were sown in sandy soil in plastic pots (240 mm length \u0026times; 240 mm width \u0026times; 190 mm depth). Pots were kept well-watered until the plants had five unfolded trifoliate leaves. Drought stress was imposed by withdrawing water for 1 day and 5 days, with soil moisture calculated by gravimetric water content. Untreated plants were used as controls. Shoots and roots were collected and immediately frozen in liquid nitrogen for RNA extraction. Total RNA was isolated from shoots and roots using TRIzol reagent following the manufacturer\u0026rsquo;s protocol. Subsequently, 5 \u0026micro;g of total RNA was mixed with 500 ng of oligo(dT)\u003csub\u003e18\u003c/sub\u003e and 200 U Superscript\u0026trade; III reverse transcriptase (Invitrogen), and the mixture was reverse transcribed at 50\u0026deg;C for 60 min. Real-time PCR was performed using the Luna\u0026reg; Universal qPCR Master Mix (NEB) with the gene-specific primers according to the manufacturer\u0026rsquo;s protocol. Relative gene expression ratios were calculated using the 2\u003csup\u003e\u0026minus;ΔΔCT\u003c/sup\u003e method, with untreated plants cDNA as the reference sample and actin as the reference gene. All experiments were conducted in three biological replicates. Statistical analysis to validate significant differences in key DTgenes transcript quantities under drought stress was performed using one-way ANOVA followed by Tukey\u0026rsquo;s HSD test method facilitated by IBM SPSS statistics software. \u003cem\u003ep\u003c/em\u003e-values less than 0.05 were considered statistically significant.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eThe development of a systematic data-driven pipeline for feature engineering (DPFE) to identify key drought-tolerant genes (DTgenes) is a critical step for the advancement of precision breeding (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). This comprehensive pipeline consists of several layers designed to extract and prioritize important genetic information from diverse data sources, including the data source layer, data processing layer, data integration layer, data prioritization layer, data systems layer, and evaluation layer. The pipeline performs thorough data filtering, quality control, and preservation of biologically valuable data across layers, culminating in the identification of key DTgenes from a large-scale feature pool (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eData acquisition and curation\u003c/h2\u003e \u003cp\u003eA total of 788,696 raw data entries were mined from cloud-based data sources, characterized by their diverse data types and formats (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e; Supplementary data 1). These data comprised: (1) Structured data: 781,492 entries stored exclusively in spreadsheet format. (2) Semi-structured data: 3,034 entries in various formats, including PDFs with extractable tables (1,289 entries), CSV (1,695 entries), and HTML. (3) Quasi-structured data: 55 entries in PDFs with unextractable tables. (4) Unstructured data: 3,665 entries in image format. Additionally, we extracted 4,896 biological pathways along with 48,606 unique soybean genes from the GO website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://geneontology.org/\u003c/span\u003e\u003cspan address=\"http://geneontology.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eData ETL and Data standardization\u003c/h2\u003e \u003cp\u003eData ETL and standardization processes were applied to ensure the consistency and usability of diverse datasets obtained from high-throughput technologies (omics data) and traditional molecular methods (non-omics data). The omics datasets included genomics, transcriptomics, proteomics, and metabolomics, while the non-omics data comprised genomic, transcriptomic, proteomic, metabolomic, and epigenomic data categories. Consequently, five OnO data platforms were established: genomic, transcriptomic, proteomic, metabolomic, and epigenomic data platforms.\u003c/p\u003e \u003cp\u003eFrom structured data formats, a total of 139,428 genomic data points, 577,788 transcriptomic data points, 60,112 proteomic data points, and 4,614 metabolomic data points were extracted. Semi-structured data contributed an additional 1,037 genomic data points, 1,748 transcriptomic data points, 199 proteomic data points, and 50 metabolomic data points. Quasi-structured data yielded 55 genomic data points, while unstructured data provided 3,647 transcriptomic data points and 9 metabolomic and epigenomic data points. The detailed hierarchical information of data types, formats, and OnO data platforms is illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eData integration and systematic omics swap\u003c/h2\u003e \u003cp\u003eA total of 147,590 data points were processed and integrated across five OnO data platforms, encompassing 49,836 genomic, 46,713 transcriptomic, 47,931 proteomic, 3,080 metabolomic, and 30 epigenomic data points. To quantify the association with drought-tolerant responses, we adopted genetic scores (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e) across these platforms and calculated the important index for individual features (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e). The genetic scores before rescaling exhibited diverse ranges: genomic (0.000007 to 9.78), transcriptomic (0.000000038 to 30.78), proteomic (1 to 16.46), metabolomic (1.33 to 8.9), and epigenomic (1.41 to 2.2).\u003c/p\u003e \u003cp\u003eThe score distribution for each OnO data platform showed right skewness and was subsequently rescaled into the range of [0, 6], except for epigenomic data platform. Summing up the rescaled \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e across the five OnO data platforms resulted in 55,465 unique features in the training gene set. The ranges of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{j}^{DR}\\)\u003c/span\u003e\u003c/span\u003e were [0.7, 2.0] and [0.000011, 29.11], respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo ensure the robustness and comprehensiveness of our pipeline, we conducted a systematic omics swap involving one genomics database and two transcriptomics databases. Following the same procedure, we obtained 106,593 data points, including 3,098 genomic and 52,454 transcriptomic data points. The number of data points remained unchanged in the other OnO data platforms. This process resulted in 54,086 unique features in the training gene set, with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e ranging from 0.7 to 1.8 and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{j}^{DR}\\:\\)\u003c/span\u003e\u003c/span\u003eranging from 0.00000044 to 26.15 (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eData prioritization\u003c/h2\u003e \u003cp\u003eThe objective of feature prioritization is to retain the most critical biomarkers for subsequent analyses. To achieve this, we applied a non-parametric random forest (NPRF) algorithm, utilizing gene frequency computation to select features. This process effectively reduced the initial pool of over 55,000 features to a refined set of 135 features (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003ea). These features, considered as biomarkers, demonstrated significant differences in gene expression levels within the validation gene set (p\u0026thinsp;\u0026lt;\u0026thinsp;0.00001). Applying the same procedure to the omics swap scenario resulted in the selection of 133 features (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eb), which were also identified as biomarkers. These biomarkers showed statistical significance within the validation gene set (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05) (Supplementary data 2).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eIntegrative systems biology analysis\u003c/h2\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003ePathways enrichment analysis\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eIn this study, we first applied a hypergeometric test for pathway enrichment to reveal critical pathways involved in drought-tolerant responses. We identified 21 (omics unswapped scenario) and 24 (omics swap scenario) GO pathways that were significantly associated with drought-tolerant responses based on 135 and 133 biomarkers, respectively. A total of 112 biomarkers were identified in enriched pathways in both scenarios, with 60 overlapping biomarkers, highlighting the complexity of the underlying biological mechanisms.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eGene network analysis\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eTo elucidate the complex nature of drought tolerance, we utilized the built-in tools within SoyNet to explore multiple pathways and gene interactions between selected biomarkers. We set up two strategies to identify functional modules. First, we constructed two distinct functional gene networks using 135 biomarkers (from omics unswapped scenario) and 133 biomarkers (from omics swap scenario) identified in feature prioritization, comprised 162 nodes and 1,778 edges, and 139 nodes and 1,268 edges, respectively. Second, we selected 112 biomarkers identified in enriched pathways from each scenario for gene network construction, forming complex gene networks comprising 136 nodes and 1,391 edges (from omics unswapped scenario), and 139 nodes and 1,393 edges (from omics swap scenario), respectively. To construct an integrative gene network, we utilized Cytoscape [42] to pool two gene networks constructed by distinct strategies for each scenario, resulting in comprehensive functional interactions comprising 162 nodes and 1,717 edges in the omics unswapped scenario (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ea) and 160 nodes and 1,393 edges in the omics swap scenario (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eTopology analysis\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe conducted a topology analysis to identify genes with significant degrees of interaction to reveal their potential role in drought tolerance. Degrees of functional interactions for each of two scenarios ranged from 0 to 105 (omics unswapped) and 0 to 97 (omics swap), respectively. We further identified 52 (out of 162 genes) and 65 (out of 160 genes) biomarkers within the 40th to 80th percentile degree numbers between 3\u0026ndash;53 (omics unswapped) and 2\u0026ndash;45 (omics swap), respectively. Among them, 26 biomarkers overlapped. After eliminating genes with zero-degree numbers, we identified 47 biomarkers (35 biomarkers from the omics unswapped scenario and 21 biomarkers from the omics swap scenario, with 9 overlapped biomarkers) and 53 intermediate genes (33 genes from the omics unswapped scenario and 31 genes from the omics swap scenario, with 11 overlapped genes) that lay outside this range.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eKey DTgenes discovery\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe identified 26 biomarkers from the intersection of two gene networks (omics unswapped network and omics swap network) within the 40th and 80th percentile of degree numbers as DTgenes candidates. We applied three selection criteria based on the distribution of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{S}_{ij}^{G}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e (please refer to the \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003eMaterials and Methods\u003c/span\u003e section for details) for further refine these candidates. Consequently, 10 candidate DTgenes were identified and regarded as key DTgenes in this study (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eEvaluations\u003c/h2\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eComputational methods\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e(1) Statistical validation\u003c/p\u003e \u003cp\u003eThe 10 key DTgenes in the corresponding gene networks statistically outperformed (p-values\u0026thinsp;\u0026lt;\u0026thinsp;0.05) the other biomarkers within gene networks (i.e., other DTgenes), intermediate genes, and the remaining genes across all drought treatment spans (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e(2) Previous research evidence\u003c/p\u003e \u003cp\u003eMultiple lines of evidence were identified across various biological levels (Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e). Only one key DTgene, \u003cem\u003eGlyma.05g056300\u003c/em\u003e, has supporting evidence at the RNA level. At the protein level, four key DTgenes (\u003cem\u003eGlyma.09g178100\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400\u003c/em\u003e, and \u003cem\u003eGlyma.19g106800\u003c/em\u003e) were associated with protein abundance research in response to drought treatment. Their encoded protein expression has been proven to be upregulated during drought treatment. At the functional level, four key DTgenes (\u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400\u003c/em\u003e, and \u003cem\u003eGlyma.19g106800\u003c/em\u003e) were identified within the Plant Metabolic Network (PMN) database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://plantcyc.org/\u003c/span\u003e\u003cspan address=\"https://plantcyc.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and were associated with five shared metabolic pathways, including Glycolysis I (from glucose-6-phosphate), Glycolysis II (from fructose-6-phosphate), Glycolysis IV (plant cytosol), Calvin-Benson-Bassham cycle, and sucrose biosynthesis I (from photosynthesis). These metabolic pathways are related to glycolysis, photosynthesis, and carbohydrates metabolism. Six key DTgenes demonstrated evidence at the homologs level (\u003cem\u003eGlyma.05g056300\u003c/em\u003e, \u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.07g043600\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400\u003c/em\u003e, and \u003cem\u003eGlyma.19g106800\u003c/em\u003e). Among them, \u003cem\u003eGlyma.07g043600\u003c/em\u003e and \u003cem\u003eGlyma.19g106800\u003c/em\u003e lacked evidence in legume crops and Arabidopsis, respectively. All top hits evidence of legume crops were found in \u003cem\u003ePisum sativum\u003c/em\u003e (pea) with query cover and percent identity over 99% and 86%, respectively. Five key DTgenes were found to have evidence in Arabidopsis through gene editing for homologs level (\u003cem\u003eGlyma.05g056300\u003c/em\u003e, \u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400\u003c/em\u003e, and \u003cem\u003eGlyma.16g211700\u003c/em\u003e). These were proved to affect plant growth under drought conditions through gene knockout techniques in Arabidopsis (\u003cem\u003eGlyma.05g056300\u003c/em\u003e) or to lose crucial physiological mechanisms associated with drought tolerance response (\u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400\u003c/em\u003e, and \u003cem\u003eGlyma.16g211700\u003c/em\u003e). No evidence was observed in soybean and model plants for two key DTgenes, \u003cem\u003eGlyma.09g163900\u003c/em\u003e and \u003cem\u003eGlyma.13g341500\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eMolecular experiments\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWe performed qRT-PCR to examine the expression of 10 key DTgenes under drought-induced stress to verify the findings of the transcriptome data of DTgenes (Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e). In the current study, the expression of five genes (\u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.07g043600\u003c/em\u003e, \u003cem\u003eGlyma.09g163900\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, and \u003cem\u003eGlyma.12g037400\u003c/em\u003e) in soybean shoot and root was considerably upregulated after withdrawing water for 5 days. The \u003cem\u003eGlyma.05g056300\u003c/em\u003e in the shoot was upregulated at 1 day and downregulated at 5 days, while in the root was significantly upregulated from 1 to 5 days. The highest expression of \u003cem\u003eGlyma.09g178100\u003c/em\u003e in the shoot was found at 5 days and in the root at 1 day. A pattern of downregulation was observed for \u003cem\u003eGlyma.13g341500\u003c/em\u003e. One gene (\u003cem\u003eGlyma.16g211700\u003c/em\u003e) was induced at 1 day, then slightly down at 5 days. Finally, no significant difference was observed in the shoot for \u003cem\u003eGlyma.19g106800\u003c/em\u003e, but the highest level was found at 5 days in the root system.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eDrought presents a significant challenge in agriculture, consistently leading to the most substantial losses among various abiotic stresses in crop yields and economic impacts globally. In 2019, drought events resulted in a 7% reduction in soybean production, translating to over 12\u0026nbsp;billion US dollars in global losses [2, 45]. Given the crucial role of soybean in food, agriculture, industry, and health sectors, these figures highlight the urgent need to develop strategies to mitigate the impacts of drought. Conventional methods for identifying important biomarkers are time-consuming and labor/resource-intensive [46\u0026ndash;48] and often unable to keep pace with the rapid environmental changes induced by climate change. Moreover, the multifaceted nature of drought-tolerant responses, involving complex interactions between various genetic and phenotypic traits [49], requires a more holistic approach.\u003c/p\u003e \u003cp\u003eIn this study, we proposed that integrating diverse biological data using a comprehensive DPFE can enhance the discovery of key DTgenes in soybean. This method ensures that only biologically valuable data are preserved and analyzed, increasing the likelihood of identifying robust DTgenes. Despite potential challenges such as data heterogeneity and computational complexity, the comprehensive design of pipeline and rigorous quality control on data enhance its effectiveness. By integrating structured, semi-structured, quasi-structured, and unstructured data, we ensured comprehensive coverage of relevant information, while advanced feature extraction techniques enhance the reliability of the findings. The subsequent data standardization and mapping techniques facilitated the alignment of heterogeneous data sources, crucial for meaningful analysis. The development of DPFE aims to not only reduce the cost and timeline for plant breeding but also enhance precision and applicability of key genes identification. By using resources from previous research (multi-omics and non-omics datasets), data-driven techniques, and systems biology, our pipeline streamlines the process of key DTgenes selection, ensuring an effective approach to enhance drought tolerance in soybean crops. This approach addresses current limitations in single-study analyses and sets a precedent for future research in crop improvement under changing climatic conditions.\u003c/p\u003e \u003cp\u003eWe conducted extensive data mining and curation to deal with data heterogeneity (distinct data formats and types) and data quality (presence of false positives (FPs) and noise, data incompleteness, and irrelevant data), ending up with 788,696 genetic raw data. Managing large amounts of data remains challenging due to the presence of FPs and noise. To address these problems, we developed tailored strategies for information retrieval from each data type and conducted data quality assurance at each layer to ensure data accuracy and reliability. By acquiring and preprocessing datasets specifically related to drought responses and organizing them based on their structural characteristics, we streamlined the subsequent data processing and analysis steps, leading to more accurate and meaningful insights into soybean drought tolerance mechanisms.\u003c/p\u003e \u003cp\u003eThe data ETL process is crucial in deriving meaningful biological information from multiple data types and formats. The variation in data formats introduced complexity into data ETL process, increasing the difficulty and time required for data standardization. This process ensures that data from diverse sources can be reliably compared and analyzed, leading to more accurate and meaningful biological insights. Raw data extracted from quasi-structured PDFs and images required more time and manual efforts in data processing to convert and standardize the data to uniform formats [50]. Integration of non-structured and structured data is a significant task in the ETL process to ensure data integrity. Continuous monitoring, validation, and refinement of the ETL process are crucial.\u003c/p\u003e \u003cp\u003eThe integration of data from multiple data platforms and multidisciplinary research was crucial to minimize the underlying risks of FPs, bias, and noise. Unlike traditional approaches that may rely on single-omics or solitary molecular technique, our approach incorporates feature selection with a conceptual integration strategy across five OnO data platforms. Numerous studies have demonstrated that a comprehensive model containing multiple OnO data can compensate for uncertain or missing data in any unbalanced dataset, outperforming models including only single-omics alone [51, 52]. This holistic approach is based on technological advancements and large amounts of previous research, revolutionizing our understanding of soybean genetics for agronomic traits. Although non-omics data were more inconsistent and variable, posing challenges for data integration [52], a well-designed conceptual integration strategy can cope with underlying FPs and compensate for the noise and bias of each data platform. The conceptual integration strategy allowed a more flexible structure to determine the mathematical approaches for fitting the existing dataset\u0026rsquo;s patterns and research objectives compared to the statistical and model-based integration strategy [37], enhancing the potential for advancing our understanding of drought-tolerant responses in soybean.\u003c/p\u003e \u003cp\u003eTo effectively integrate and analyze this multifaceted data, we developed a scoring scheme, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e, and normalization process. By rescaling the score of data points within different data platforms into a unified numerical score [0, 6], we ensured a common scale across entire OnO data platforms, addressing the inherent variability in molecular properties. In terms of the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{II}_{j}\\)\u003c/span\u003e\u003c/span\u003e index, which increases the sensitivity for biomarker identification, by increasing (or decreasing) the weight of features with strong or consistent (or weak or inconsistent) evidence, the system ensures that only the most credible candidates are highlighted as key DTgenes in the subsequent analysis. Through this strategy, we can improve the high FP rate of biomarkers findings in traditional breeding progress. The systematic omics swap involved a genomics database and two transcriptomics databases. The primary reason for selecting the genomics database as the swapped database was its highly comprehensive genome data, containing over 130,000 SNPs mapped to over 44,000 soybean reference genes. This approach produced an alternative gene pool, enhancing the comprehensiveness of features for the identifying biomarkers more likely to confer drought tolerance in soybean.\u003c/p\u003e \u003cp\u003eThe application of the NPRF algorithm for feature selection was pivotal in the current study. We strategically determined the number of features by the distribution of gene frequency, setting the selection threshold just before the first significant drop in frequency. The feature selection process reduced the feature pool from over 55,000 to 135 features, ensuring the meaningful and manageable biomarker candidates. The combination of ETL process, conceptual integration, and feature selection facilitated fast and precise biomarker discovery.\u003c/p\u003e \u003cp\u003eThe integrative approach combining multiple omics datasets and systematic omics swap provides a powerful framework for robust biomarker discovery. The consistency observed in the omics swap scenario further validates the reliability of our feature selection method. The rationale behind the omics swap procedure is to mitigate the risk of overfitting and ensure that the identified biomarkers are not specific to a particular subset of the data. This method allows for cross-validation of key genes and reduces the risk of FPs, enhancing the reliability of our findings. The insights gained from this integrative analysis can significantly contribute to advancing precision breeding strategies aimed at improving drought tolerance in soybeans.\u003c/p\u003e \u003cp\u003eThe surge of biological data has fundamentally shifted the paradigm of systems biology in plants [53]. Through integrative pathway enrichment and construction of gene networks, we have successfully obtained novel insights into the molecular interactions and functional implications of complex regulatory networks governing drought-tolerant responses in soybean. Gene network and topological analysis emerged as complex yet informative, highlighting critical nodes and connections key to the adaptation to drought in soybean. Biomarkers involved with common biological processes are likely to be functionally related. Genes with higher connectivity (degree number), however, do not necessarily imply greater functional relevance. Genes with intermediate connectivity are considered to have the highest functional co-annotation rate [54, 55]. To find the most functionally related biomarkers, we eliminated biomarkers with a degree number ranking below the 40th percentile, as they contribute minimally to the network. Conversely, we also excluded biomarkers with a degree number above the 80th percentile to avoid overly general biomarkers with lower specific functional relationships. This method allowed us to focus on biomarkers most likely integral to drought tolerance in soybean.\u003c/p\u003e \u003cp\u003eBy combining two gene networks, 10 key DTgenes were identified. These key DTgenes potentially represent the pivotal role in the soybean drought-tolerant responses, supported by high functional relationship within gene networks and enriched pathways. However, no key DTgenes were identified in the epigenomic data platform due to limited relevant information. Nine out of ten key DTgenes (\u003cem\u003eGlyma.05g056300\u003c/em\u003e, \u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.07g043600, Glyma.09g163900\u003c/em\u003e, \u003cem\u003eGlyma.09g178100\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400, Glyma.13g341500\u003c/em\u003e, \u003cem\u003eGlyma.16g211700\u003c/em\u003e, \u003cem\u003eGlyma.19g106800\u003c/em\u003e) were recorded across four OnO data platforms, except for \u003cem\u003eGlyma.07g043600\u003c/em\u003e, which was recorded in three data platforms (genomic, transcriptomic, and proteomic). Of these key DTgenes, the highest frequency was identified in \u003cem\u003eGlyma.06g015900\u003c/em\u003e, indicating its involvement in protein-protein interactions and biological processes. Despite \u003cem\u003eGlyma.07g043600\u003c/em\u003e being recorded in only three platforms, it showed significant scores in genomic, transcriptomic, and proteomic data, suggesting its importance in drought stress response.\u003c/p\u003e \u003cp\u003eThe integration of large-scale multiple OnO data has provided a valuable resource for identifying potential key DTgenes. Nevertheless, the reliability and functionality of these genes must be rigorously assessed through further research. We expect that these genes will exhibit significant changes across various biological levels. Legume crops were selected as model plant species due to their close phylogenetic relationship, facilitating the inference of functions for key DTgenes. Not all key DTgenes, however, have identifiable homologous genes or functions in legume crops. To bridge this gap, Arabidopsis was introduced as a complementary model species due to its well-annotated functional genome. Sequence identity in BLAST is the widely-known strategy for identification of homologous proteins and genes [56]. We searched for corresponding Arabidopsis homologs in SoyBase based on the highest sequence similarity associated with key DTgenes. In this study, we successfully identified ten key DTgenes in soybean using a data-driven pipeline, with eight having corresponding evidence in previous research and two (\u003cem\u003eGlyma.09g163900\u003c/em\u003e and \u003cem\u003eGlyma.13g341500\u003c/em\u003e) lacking relevant supporting documents.\u003c/p\u003e \u003cp\u003eAmong the identified key DTgenes, \u003cem\u003eGlyma.05g056300\u003c/em\u003e emerged as a prime candidate, showing strong evidence across multiple levels, including RNA, homologs, and gene editing in Arabidopsis. This gene encodes acid-soluble invertase in soybean (GmA-INV / GmA-VINV1.1A) and exhibits significant expression changes during drought at key flowering stages, suggesting a pivotal role in drought response [57]. Its orthologs in pea (\u003cem\u003ePisum sativum\u003c/em\u003e) [58] (\u003cem\u003ePsVINV1.1\u003c/em\u003e) and Arabidopsis [59] (\u003cem\u003eAt1g12240\u003c/em\u003e /\u003cem\u003eAtVIN2\u003c/em\u003e) have shown similar regulation, indicating drought resistance. Stomatal conductance is significantly reduced in \u003cem\u003eVIN2\u003c/em\u003e-suppressed Arabidopsis plants, which are less tolerant to water deficit than wild-type plants, reinforcing its importance [60].\u003c/p\u003e \u003cp\u003e \u003cem\u003eGlyma.11g111400\u003c/em\u003e and \u003cem\u003eGlyma.12g037400\u003c/em\u003e, encoding fructose biphosphate aldolase 2 (GmFBA2), also demonstrated substantial evidence across four biological levels. There genes are homologous to an identical Arabidopsis gene (\u003cem\u003eAt4g38970\u003c/em\u003e) and are involved in critical metabolic pathways, especially under drought stress conditions in soybean cotyledons [61]. The Arabidopsis homolog (\u003cem\u003eAtFBA2\u003c/em\u003e) plays a central role in photosynthesis and energy metabolism [62] and is essential for starch synthesis in the Calvin cycle [63]. The \u003cem\u003efba2\u003c/em\u003e mutant shows growth retardation due to inhibited sugar synthesis [63]. Modifications in carbon metabolism demonstrate drought improvement in resistant cultivars, storing more carbon for energy generation to maintain physiological mechanisms under drought [64]. Additionally, the Arabidopsis homologs \u003cem\u003eAtFBA2\u003c/em\u003e (\u003cem\u003eKIW84_074187\u003c/em\u003e and \u003cem\u003eKIW84_074188\u003c/em\u003e) were observed in peas. Their encoded proteins are significantly upregulated under drought conditions, playing crucial roles in glycolytic process in pea seedlings [65].\u003c/p\u003e \u003cp\u003e \u003cem\u003eGlyma.06g015900\u003c/em\u003e, encoding glyceraldehyde-3-phosphate dehydrogenase B subunit (GAPB), is homologous to \u003cem\u003eAt1g42970\u003c/em\u003e (\u003cem\u003eAtGAPB\u003c/em\u003e). \u003cem\u003eAtGAPB\u003c/em\u003e shows significant expression changes under different drought durations [66]. The \u003cem\u003egapb\u003c/em\u003e mutant transgenic line demonstrates a significant decrease in CO\u003csub\u003e2\u003c/sub\u003e assimilation in photosynthesis, implying GAPB\u0026rsquo;s potential importance in regulating drought tolerance [67, 68]. Orthologous genes of \u003cem\u003eGlyma.06g015900\u003c/em\u003e were identified in pea, showing significant changes under water stress and osmotic stress recovery stages [69].\u003c/p\u003e \u003cp\u003eOther key DTgenes such as \u003cem\u003eGlyma.09g178100\u003c/em\u003e and \u003cem\u003eGlyma.19g106800\u003c/em\u003e, showed strong evidence in soybean [61, 70], while no other evidence was found in homologs for \u003cem\u003eGlyma.09g178100\u003c/em\u003e. The orthologs of \u003cem\u003eGlyma.19g106800\u003c/em\u003e in pea encodes glyceraldehyde-3-phosphate dehydrogenase A and shows significant protein abundance changes under water stress [69]. In previous transcriptomic research, \u003cem\u003eGlyma.19g106800\u003c/em\u003e was one of the top 50 most highly expressed transcripts under drought stress [71]. Both \u003cem\u003eGlyma.07g043600\u003c/em\u003e and \u003cem\u003eGlyma.16g211700\u003c/em\u003e were supported by homologous genes in Arabidopsis [70, 72]. The homolog of \u003cem\u003eGlyma.07g043600\u003c/em\u003e in Arabidopsis (\u003cem\u003eAt1g54050\u003c/em\u003e) is indirectly associated with drought tolerance, with significant gene expression changes under oxidative stress. Oxidative stress is induced by multiple abiotic and biotic stresses, including drought [72]. \u003cem\u003eGlyma.16g211700\u003c/em\u003e is homologous to \u003cem\u003eAt1g73260\u003c/em\u003e (Kunitz-type trypsin inhibitor 1, \u003cem\u003eAtKTI1\u003c/em\u003e) in Arabidopsis, which causes early flowering when overexpressed [73], a behavior associated with drought escape [74]. Previous proteomics studies indicate that proteins related to energy and photosynthesis play crucial roles in drought tolerance [70], aligning with our functional level discoveries.\u003c/p\u003e \u003cp\u003eFour key DTgenes were found to be involved in glycolysis, carbohydrates metabolism, and the Calvin cycle. The remaining key DTgenes, \u003cem\u003eGlyma.09g163900\u003c/em\u003e and \u003cem\u003eGlyma.13g341500\u003c/em\u003e, represent novel discoveries. While they lacked prior research evidence, their significant expression level differences across three OnO data platforms highlight their potential as new contributors to drought tolerance in soybean. These two genes are innovative discoveries made possible by our comprehensive, multi-faceted analytical approach.\u003c/p\u003e \u003cp\u003eFurther analysis using ShinyGO v0.80 [75] identified six genes (\u003cem\u003eGlyma.05g056300\u003c/em\u003e, \u003cem\u003eGlyma.06g015900\u003c/em\u003e, \u003cem\u003eGlyma.09g178100\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, \u003cem\u003eGlyma.12g037400\u003c/em\u003e, and \u003cem\u003eGlyma.19g106800\u003c/em\u003e) enriched in eight GO terms related to carbohydrate metabolism. Two genes (\u003cem\u003eGlyma.05g056300\u003c/em\u003e and \u003cem\u003eGlyma.09g178100\u003c/em\u003e) were significantly enriched in five KEGG pathways, mainly related to carbohydrate metabolism and secondary metabolite synthesis. Four KEGG pathway networks demonstrated the roles of \u003cem\u003eGlyma.05g056300\u003c/em\u003e and \u003cem\u003eGlyma.09g178100\u003c/em\u003e in these networks (Fig.\u0026nbsp;\u003cspan refid=\"Fig12\" class=\"InternalRef\"\u003e12\u003c/span\u003e). These findings suggest that improving or optimizing carbohydrate metabolism in soybean may confer drought tolerance to soybean seedlings.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe DPFE framework used in the current study, which combines data ETL and feature selection, not only identified key genes consistent with previous research but also uncovered novel genes that may play crucial roles in drought-tolerant responses. To further confirm the functionality of these genes, future work should investigate their association with drought tolerance through RNA silencing or gene knockout.\u003c/p\u003e \u003cp\u003eEnvironmental stress, including drought, significantly restricts plant growth and development. Increasing crop tolerance to drought is an effective strategy for maintaining yields [76]. In this study, most key DTgenes showed increased expression, while one gene exhibited decreased expression following drought stress, according to qRT-PCR data. Numerous plant mechanisms for environmental stress, including our earlier research [44], have demonstrated such up-and-down regulation. Three genes (\u003cem\u003eGlyma.05g056300\u003c/em\u003e, \u003cem\u003eGlyma.11g111400\u003c/em\u003e, and \u003cem\u003eGlyma.12g037400\u003c/em\u003e) related to carbohydrate transport and metabolism [77\u0026ndash;80] were upregulated, indicating that soybeans, like other plants, respond to drought stress through carbohydrate metabolism.\u003c/p\u003e \u003cp\u003eAmong the 10 key DTgenes, \u003cem\u003eGlyma.07g043600\u003c/em\u003e exhibited the highest relative gene expression. Predicted as GmHSP20, this gene was highly induced in RNA-seq data [81]. Plant heat shock proteins (Hsps) typically regulate antioxidant enzymes positively, improving membrane integrity and detoxifying reactive oxygen species during stress tolerance [82]. \u003cem\u003eGlyma.09g163900\u003c/em\u003e and \u003cem\u003eGlyma.13g341500\u003c/em\u003e, which lacked previous evidence, showed significantly different expression patterns under drought treatment in this study, suggesting they may be novel biomarkers for drought tolerance research in soybean. Thus, our qRT-PCR data may indicate the potential biological roles of key DTgenes in plant protection against drought stress.\u003c/p\u003e \u003cp\u003eDespite the promising potential of our data-driven pipeline, several challenges in the present study must be acknowledged. Firstly, our findings relied entirely on existing research, and the completeness and quality of these data could introduce noise and biases. To address this, we implemented a data ETL process with stringent standards for data quality control and management to minimize these risks. Integrating multiple data platforms based on different molecular methodologies can minimize the publication bias and FPs, reducing the risk of biased results. Secondly, the heterogeneity of data sources poses significant integration difficulties. Our data-driven pipeline employs several strategies, including data ETL and standardization, multi-platform integration, data mapping and alignment, advanced computational tools, and systematic omics swap, to address these issues effectively. Thirdly, the availability and accessibility of relevant datasets can be challenging, as some datasets may be dispersed across different repositories and require considerable effort to locate and retrieve. Our pipeline incorporates strategies for data aggregation and retrieval, data integration and standardization, and centralized repository setup to manage this challenge effectively. Finally, the applicability of this pipeline may be limited during the data integration stages due to the scarcity of experimental data in a domain. As the domain develops and more data accumulates, the discovery of biomarkers through this pipeline will become more consistent and stable, further supporting its reliability, robustness, and precision. Additionally, processing and analyzing large-scale datasets require robust infrastructure and sophisticated algorithms.\u003c/p\u003e \u003cp\u003eThe data-driven pipeline for feature engineering presented in this study represents a significant advancement in the field of crop genetics and breeding. By leveraging the power of big data and systems biology, we identified key DTgenes that hold promise for improving drought tolerance in soybean. This approach addresses the limitations of traditional breeding methods and sets the stage for future research and applications in various agricultural contexts.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e1. Ethics approval and consent to participate\u003c/p\u003e\n\u003cp\u003eThis article does not contain any studies with human participants or animals performed by any of the authors.\u003c/p\u003e\n\u003cp\u003e2. Consent for publication\u003c/p\u003e\n\u003cp\u003eNot applicable. This article does not contain any studies with human participants or animals performed by any of the authors that require consent for publication.\u003c/p\u003e\n\u003cp\u003e3. Availability of data and materials\u003c/p\u003e\n\u003cp\u003eThe raw data acquired from source of article underlying this article are available in Github (\u003cu\u003egithub.com/edwardkao6413/Raw-data-of-soybean-drought-tolerant-research.git\u003c/u\u003e), further inquiries or any data support can be directed to the corresponding author.\u003c/p\u003e\n\u003cp\u003e4. Competing interests\u003c/p\u003e\n\u003cp\u003eThe authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.\u003c/p\u003e\n\u003cp\u003e5. Funding\u003c/p\u003e\n\u003cp\u003eThis study is funded by Chung Cheng Agriculture Science and Social Welfare Foundation, FUTURETECH ELECTRONICS CO., LTD, National Science and Technology Council (NSTC 112-2622-B-005-001), and Advanced Plant and Food Crop Biotechnology Center from The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education in Taiwan.\u003c/p\u003e\n\u003cp\u003e6. Authors\u0026apos; contributions\u003c/p\u003e\n\u003cp\u003eC.F.Kao conceived the study conception and design. P.H.Kao, C.W.Li, C.W.Tseng, and C.F.Kao collected the data. P.H.Kao and S.Baiya performed the data analysis. P.H.Kao, S.Baiya, C.F.Kao, S.Y.Chen, and Y.H.Huang interpreted the data and results. P.H.Kao, S.Baiya, and C.F.Kao drafted the manuscript. C.F.Kao and P.H.Kao revised the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e7. Acknowledgements\u003c/p\u003e\n\u003cp\u003eThis study was granted by Chung Cheng Agriculture Science and Social Welfare Foundation and FUTURETECH ELECTRONICS CO., LTD. We extend our appreciation to the National Science and Technology Council (NSTC 112-2622-B-005-001). For their financial support, which contributed (in part) to the success of this research. Furthermore, we would like to thank the Advanced Plant and Food Crop Biotechnology Center from The Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education in Taiwan. Their support has been invaluable in advancing agricultural research and promoting scientific innovation.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eHodges T, French V. Soyphen: Soybean Growth Stages Modeled from Temperature, Daylength, and Water Availability. J Agron. 1985;77:500\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRazzaq A, Wani SH, Saleem F, Yu M, Zhou M. Rewilding crops for climate resilience: economic analysis and de novo domestication strategies. J Exp Bot. 2021;72:6123\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKramer PJ. J. S. Boyer. Water relations of plants and soils. San Diego: Academic; 1995.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeleiman MF, Al-Suhaibani N, Ali N, Akmal M. Alotaibi. Drought stress impacts on plants and different approaches to alleviate its adverse effects. Plants. 2021;10:259.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDesclaux D, Roumet P. Impact of drought stress on the phenology of two soybean (Glycine max L. Merr) cultivars. Field Crops Res. 1996;46:61\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLafitte H, Yongsheng G, Yan S, Li Z. Whole plant responses, key processes, and adaptation to drought stress: the case of rice. J Exp Bot. 2007;58:169\u0026ndash;75.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYadav RS, Hash C, Bidinger F, Devos K. Howarth. Genomic regions associated with grain yield and aspects of post-flowering drought tolerance in pearl millet across stress environments and tester background. Euphytica. 2004;136:265\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eByrum J, Davis C, Doonan G, Doubler T. D. Foster. Advanced analytics for agricultural product development. Interfaces. 2016;46:5\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKusmec A, Zheng Z, Archontoulis S, Ganapathysubramanian B, Hu G. Interdisciplinary strategies to enable data-driven plant breeding in a changing climate. One Earth. 2021;4:372\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmar S, Gill RA, Jung K-H, Faheem A, Qasim MU. Conventional and molecular techniques from simple breeding to speed breeding in crop plants: recent advances and future outlook. Int J Mol Sci. 2020;21:2590.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSharma N, Raman H, Wheeler D, Kalenahalli Y, Sharma R. Data-driven approaches to improve water-use efficiency and drought resistance in crop plants. Plant Sci. 2023;336:111852.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaleem A, Rold\u0026aacute;n-Ruiz I, Aper J, Muylle H. Genetic control of tolerance to drought stress in soybean. BMC Plant Biol. 2022;22:1\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeckwerth W, Ghatak A, Bellaire A, Chaturvedi P. Varshney. PANOMICS meets germplasm. Plant Biotechnol J. 2020;18:1507\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWard RM, Schmieder R, Highnam G, Mittelman D. Big data challenges and opportunities in high-throughput sequencing. Syst Biomed. 2013;1:29\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNekrutenko A, Taylor JJNRG. Next-generation sequencing data interpretation: enhancing reproducibility and accessibility. Nat Rev Genet. 2012;13:667\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKao P-H, Baiya S, Lai Z-Y, Huang C-M. Jhan. An advanced systems biology framework of feature engineering for cold tolerance genes discovery from integrated omics and non-omics data in soybean. Front Plant Sci. 2022;13:1019709.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBell G, Hey T. Szalay. Beyond the data deluge. Science. 2009;323:1297\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePal S, Mondal S, Das G, Khatua S. Ghosh. Big data in biology: The hope and present-day challenges in it. Gene Rep. 2020;21:100869.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIsaac NJ, van Strien AJ, August TA, de Zeeuw MP. Roy. Statistics for citizen science: extracting signals of change from noisy ecological data. Methods ecol evol. 2014;5:1052\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eR. M. Dorazio. Accounting for imperfect detection and survey bias in statistical analysis of presence-only data. Glob Ecol Biogeogr. 2014;23:1472\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFithian W, Elith J, Hastie T, Keith DA. Bias correction in species distribution models: pooling survey and collection data for multiple species. Methods ecol evol. 2015;6:424\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIsaac NJ, Jarzyna MA, Keil P, Dambly LI. Boersch-Supan. Data integration for large-scale models of species distributions. Trends Ecol Evol. 2020;35:56\u0026ndash;67.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLarose DT, Larose CD. An Introduction to Data Mining of 1. In: editor. An Introduction to Data Mining of Book. New Jersey: Wiley; 2014.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVijayarani S, Ilamathi MJ. Nithya. Preprocessing techniques for text mining-an overview. IJCSC. 2015;5:7\u0026ndash;16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSultana A. Unraveling the Data Structures of Big Data, the HDFS Architecture and Importance of Data Replication in HDFS. Int J Res Eng Technol. 2018;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDhaouadi A, Bousselmi K, Gammoudi MM, Monnet S. Hammoudi. Data warehousing process modeling from classical approaches to new trends: Main features and comparisons. Data. 2022;7:113.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZebari R, Abdulazeez A, Zeebaree D, Zebari D. Saeed. A comprehensive review of dimensionality reduction techniques for feature selection and feature extraction. J Appl Sci Technol Trends. 2020;1:56\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEl-Sappagh SHA, Hendawi AMA. El Bastawissy. A proposed model for data warehouse ETL processes. J KING SAUD UNIV-COM. 2011;23:91\u0026ndash;104.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSheth BP, Thaker VS. Plant systems biology: insights, advances and challenges. Planta. 2014;240:33\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHong W-J, Kim Y-J, Chandran AKN. Jung. Infrastructures of systems biology that facilitate functional genomic study in rice. RICE. 2019;12:1\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePazhamala LT, Kudapa H, Weckwerth W, Millar AH. Varshney. Systems biology for crop improvement. Plant Genome. 2021;14:e20098.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFukushima A, Kusano M, Redestig H, Arita M, Saito K. Integrated omics approaches in plant systems biology. Curr Opin Chem Biol. 2009;13:532\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFukushima A, Kanaya S, Nishida K. Integrated network analysis and effective tools in plant systems biology. Front Plant Sci. 2014;5:598.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYoosefzadeh Najafabadi M, Hesami M, Eskandari M. Machine learning-assisted approaches in modernized plant breeding programs. Genes. 2023;14:777.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShin JH, Vaughn JN, Abdel-Haleem H, Chavarro C. Abernathy. Transcriptomic changes due to water deficit define a general soybean response and accession-specific pathways for drought avoidance. BMC Plant Biol. 2015;15:1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTamang BG, Li S, Rajasundaram D, Lamichhane S, Fukao T. Overlapping and stress-specific transcriptomic and hormonal responses to flooding and drought in soybean. Plant J. 2021;107:100\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHahne F, Huber W, Gentleman R, Falcon S. Falcon. Hypergeometric testing used for gene set enrichment analysis of 14. In: editor. Hypergeometric testing used for gene set enrichment analysis of Book. New York: Springer; 2008. pp. 207\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim E, Hwang S, Lee I. SoyNet: a database of co-functional networks for soybean Glycine max. Nucleic Acids Res. 2017;45:D1082\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaito R, Smoot ME, Ono K, Ruscheinski J. Wang. A travel guide to Cytoscape plugins. Nat Methods. 2012;9:1069\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMutava RN, Prince SJK, Syed NH, Song L, Valliyodan B. Understanding abiotic stress tolerance mechanisms in soybean: A comparative evaluation of soybean response to drought and flooding stress. Plant Physiol Biochem. 2015;86:109\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJhan L-H, Yang C-Y, Huang C-M, Lai M-C. Huang. Integrative pathway and network analysis provide insights on flooding-tolerance genes in soybean. Sci Rep. 2023;13:1980.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim W, Iizumi T, Nishimori M. Global patterns of crop production losses associated with droughts from 1983 to 2009. J Appl Meteorol Climatol. 2019;58:1233\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu W, Wang M, Fu S, Yu D. Mapping QTLs for seed yield and drought susceptibility index in soybean (Glycine max L.) across different environments. J GENET GENOMICS. 2009;36:721\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu W, Yu D, Fu S. Detection of quantitative trait loci for yield and drought tolerance traits in soybean using a recombinant inbred line population. J Integr Plant Biol. 2009;51:868\u0026ndash;78.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSpecht J, Chase K, Macrander M, Graef G, Chung J. Soybean response to water: a QTL analysis of drought tolerance. Crop Sci. 2001;41:493\u0026ndash;509.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShelake RM, Kadam US, Kumar R, Pramanik D, Singh AK. Engineering drought and salinity tolerance traits in crops through CRISPR-mediated genome editing: Targets, tools, challenges, and perspectives. Plant Commun. 2022;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang W, Zeng Y, Jiao M, Ye C, Li Y. Integration of high-throughput omics technologies in medicinal plant research: The new era of natural drug discovery. Front Plant Sci. 2023;14:1073848.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL\u0026oacute;pez de Maturana E, Alonso L, Alarc\u0026oacute;n P, Mart\u0026iacute;n-Antoniano IA. Pineda. Challenges in the integration of omics and non-omics data. Genes. 2019;10:238.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGreener JG, Kandathil SM, Moffat L, Jones DT. A guide to machine learning for biologists. Nat Rev Mol Cell Biol. 2022;23:40\u0026ndash;55.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHou J, van Leeuwen J, Andrews BJ, Boone C. Genetic network complexity shapes background-dependent phenotypic expression. Trends Genet. 2018;34:578\u0026ndash;86.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiegal ML, Promislow DE, Bergman A. Functional and evolutionary inference in gene networks: does topology matter? Genetica. 2007;129:83\u0026ndash;103.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eW. R. Pearson. An introduction to sequence similarity (homology) searching. Curr. Protoc. Bioinform. 2013;42:3.1. 1-3.1. 8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu Y, Zhao Q, Chen L, Yao X, Zhang H. Effect of drought stress during soybean R2\u0026ndash;R6 growth stages on sucrose metabolism in leaf and seed. Int J Mol Sci. 2020;21:618.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMorin A, Kadi F, Porcheron B, Vriet C, Maurousset L. Genome-wide identification of invertases in Fabaceae, focusing on transcriptional regulation of Pisum sativum invertases in seed subjected to drought. Physiol Plant. 2022;174:e13673.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNi DA. Role of vacuolar invertase in regulating Arabidopsis stomatal opening. Acta Physiol Plant. 2012;34:2449\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S-F, Liang K, Yin D-M, Ni D-A. Zhang. Ectopic expression of a tobacco vacuolar invertase inhibitor in guard cells confers drought tolerance in Arabidopsis. J Enzyme Inhib Med. 2016;31:1381\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang X, Oh M, Sakata K, Komatsu S. Gel-free/label-free proteomic analysis of root tip of soybean over time under flooding and drought stresses. J Proteom. 2016;130:42\u0026ndash;55.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZadražnik T, Hollung K, Egge-Jacobsen W, Meglič V. Šuštar-Vozlič. Differential proteomic analysis of drought stress response in leaves of common bean (Phaseolus vulgaris L). J Proteom. 2013;78:254\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu W, Tang X, Huo Y, Xu R, Qi S. Identification and characterization of fructose 1, 6-bisphosphate aldolase genes in Arabidopsis reveal a gene family with diverse responses to abiotic stresses. Gene. 2012;503:65\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCuellar-Ortiz SM, De La Paz M, Arrieta‐Montiel J, Acosta‐Gallegos AA. Covarrubias. Relationship between carbohydrate partitioning and drought resistance in common bean. PLANT CELL ENVIRON. 2008;31:1399\u0026ndash;409.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJulia S, Daria G, Veronika P, Alexander T, Alena K. Protein glycation and drought response of pea (Pisum sativum L.) root nodule proteome: A proteomics approach. Commun Biol. 2021;66:210\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeng L, Deng R, Guo Z, Yang S, Deng X. Genome-wide identification and characterization of Glyceraldehyde-3-phosphate dehydrogenase genes family in wheat (Triticum aestivum). BMC Genom. 2016;17:1\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFeller U. Drought stress and carbon assimilation in a warming climate: Reversible and irreversible impacts. J Plant Physiol. 2016;203:84\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlqurashi M. Investigating the Role of RAP2. 12 and GAPDH in Photosynthetic CO2 Fixation to Improve Drought Tolerance [Doctoral dissertation]. Place: Institution; 2019.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrosowska-Arendt W, Gallardo K, Sommerer N, Weidner S. Changes in the proteome of pea (Pisum sativum L.) seeds germinating under optimal and osmotic stress conditions and subjected to post-stress recovery. Acta Physiol Plant. 2014;36:795\u0026ndash;807.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu X, James A, Yang A, Jones A, Mendoza-Porras O. A comparative proteomic study of drought-tolerant and drought-sensitive soybean seedlings under drought stress. Crop Pasture Sci. 2016;67:528\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen W, Yao Q, Patil GB, Agarwal G, Deshmukh RK. Identification and comparative analysis of differential gene expression in soybean leaf tissue under drought and flooding stress revealed by RNA-Seq. Front. Plant Sci. 2016;7:1044.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDat J, Vandenabeele S, Vranova E, Van Montagu M, Inz\u0026eacute;* D. Dual action of the active oxygen species during plant stress responses. Cell Mol Life Sci. 2000;57:779\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim J-g, Baek S-A, Im K-H. Overexpression of a Kunitz-type trypsin inhibitor (AtKTI1) causes early flowering in Arabidopsis. Plant Growth Regul. 2009;59:75\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoerma HR. J. E. Specht. Soybeans: improvement, production and uses. American Society of Agronomy; 2004.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGe SX, Jung D, Yao R. ShinyGO: a graphical gene-set enrichment tool for animals and plants. Bioinformatics. 2020;36:2628\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShan L, Xu Y, Wu D, Hu J. T. Yu. Effects of Salicylic Acid on Growth, Physiology, and Gene Expression in Rice Seedlings Under Salt and Drought Stress. Plant Stress. 2024;100413.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu X, Fu X, Yang Q, Jin H, Zhu L. Genome-wide variation analysis of four vegetable soybean cultivars based on re-sequencing. Plants. 2021;11:28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKumar V, Goyal V, Mandlik R, Kumawat S, Sudhakaran S. Pinpointing genomic regions and candidate genes associated with seed oil and protein content in soybean through an integrative transcriptomic and QTL meta-analysis. Cells. 2022;12:97.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSu L, Xu C, Zeng S, Su L, Joshi T. Large-Scale Integrative Analysis of Soybean Transcriptome Using an Unsupervised Autoencoder Model. Front. Plant Sci. 2022;13:831204.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchl\u0026auml;pfer P, Zhang P, Wang C, Kim T, Banf M. Genome-wide prediction of metabolic enzymes, pathways, and gene clusters in plants. Plant Physiol. 2017;173:2041\u0026ndash;59.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiao R, Han Q, Liu Y, Zhang X, Hao Q. Melatonin Attenuates the Urea-Induced Yields Improvement Through Remodeling Transcriptome and Rhizosphere Microbial Community Structure in Soybean. Front microbiol. 2022;13:903467.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLone RA, Sarvendra K, Singh V, Bano N, Bag SK. Adaptation of winged bean (Psophocarpus tetragonolobus (L.) DC.) to drought stress is mediated by root-tuber heat-shock proteins and specific metabolites. Curr Plant Biol. 2022;32:100266.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"journal-of-big-data","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bigd","sideBox":"Learn more about [Journal of Big Data](http://journalofbigdata.springeropen.com)","snPcode":"40537","submissionUrl":"https://submission.nature.com/new-submission/40537/3","title":"Journal of Big Data","twitterHandle":"@SpringerOpen","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"soybean, drought-tolerant responses, data-driven, systems biology, multidisciplinary, biomarkers discovery","lastPublishedDoi":"10.21203/rs.3.rs-4791767/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4791767/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSoybean [\u003cem\u003eGlycine max\u003c/em\u003e (L.) Merr.] is an important global food crop but is highly vulnerable to environmental changes, particularly drought. Conventional strategies of biomarkers discovery for developing drought-tolerant varieties are resource-intensive, inefficient, and without comprehensive insight. The current study proposed a novel data-driven pipeline for feature engineering through integrating diverse genetic data from multidisciplinary research on cloud-based sources in identification of key drought-tolerant genes (DTgenes) in soybean. Our pipeline involved data extraction, transformation, loading, and systematic integration of both omics and non-omics data. Feature prioritization was performed for feature selection to uncover the important biomarkers from feature pool (candidate gene pool), and key DTgenes were identified through integrative systems biology strategies based on the biomarker candidates. The validation of identified key DTgenes was carried out by both computational and molecular experiments approach. Both approaches demonstrated the credibility and potential of the key DTgenes for conferring drought tolerance response in soybean. This data-driven pipeline for feature engineering approach enhances the efficiency and accuracy of biomarkers discovery for further breeding program, which uncover a robust key DTgenes candidate which contribute to drought tolerance improvement in soybean and show potential of applicability for other crop species.\u003c/p\u003e","manuscriptTitle":"Unveiling insights into drought-tolerant responses in soybean: a data-driven pipeline for feature engineering in biomarker discovery","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-22 06:10:58","doi":"10.21203/rs.3.rs-4791767/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-12-20T13:56:19+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-12-20T13:52:19+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-12-20T13:51:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"82105712062998161324866156620992066990","date":"2024-12-18T20:55:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"26703657026904408867261167682436178684","date":"2024-12-18T20:54:51+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-12-12T18:46:45+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-12-12T08:09:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"7554900943877543389603481270300621642","date":"2024-12-04T06:39:47+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"161162671161865118770982605473612556277","date":"2024-12-04T05:52:33+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-09-02T10:50:36+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-08-16T21:10:06+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-07-26T05:00:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Big Data","date":"2024-07-24T02:21:39+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"journal-of-big-data","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bigd","sideBox":"Learn more about [Journal of Big Data](http://journalofbigdata.springeropen.com)","snPcode":"40537","submissionUrl":"https://submission.nature.com/new-submission/40537/3","title":"Journal of Big Data","twitterHandle":"@SpringerOpen","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"41a67595-04a6-4420-bd67-a4e01b489196","owner":[],"postedDate":"August 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-03-24T16:02:59+00:00","versionOfRecord":{"articleIdentity":"rs-4791767","link":"https://doi.org/10.1186/s40537-025-01078-w","journal":{"identity":"journal-of-big-data","isVorOnly":false,"title":"Journal of Big Data"},"publishedOn":"2025-03-19 15:57:59","publishedOnDateReadable":"March 19th, 2025"},"versionCreatedAt":"2024-08-22 06:10:58","video":"","vorDoi":"10.1186/s40537-025-01078-w","vorDoiUrl":"https://doi.org/10.1186/s40537-025-01078-w","workflowStages":[]},"version":"v1","identity":"rs-4791767","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4791767","identity":"rs-4791767","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.