Interpretable Machine Learning on Soybean Multi Omics Data Reveals Drought-Driven Shifts of Plant-Microbe Interactions

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-05

This study utilized interpretable machine learning on soybean multi-omics data to identify drought-driven shifts in plant-microbe interactions.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-05 · read from full text

The study integrated genomic, metabolomic, and microbiome data from 198 soybean accessions grown under control and drought conditions to predict plant phenotypes, comparing best linear unbiased prediction, GWAS, and nonlinear machine learning models for their ability to detect informative features. The authors found that machine learning models performed better than linear approaches in capturing nonlinear dependencies, using flexible variable selection, and that SHAP-based interpretation highlighted daidzin (an isoflavone derivative) and the drought-tolerant bacterium Candidatus Nitrosocosmicus as major contributors to drought-associated phenotypic variation. SHAP interaction networks further suggested cross-omics links, including associations among daidzin, GABA, and Paenibacillus. The paper is primarily about plant–microbe responses to drought and does not discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background Plant-microbe interactions in the rhizosphere are central to plant growth, nutrient acquisition, and stress resilience. Although multi-omics approaches enable comprehensive profiling of different biological layers, integrating these data to understand the mechanisms underlying plant-microbe symbiosis, particularly under drought stress, remains a challenge. Results Genomic, metabolomic, and microbiome data from 198 soybean accessions grown under both control and drought conditions were integrated to identify environment-specific predictive features of the plant phenotypes. We compared best linear unbiased prediction (BLUP), genome-wide association study (GWAS), and a nonlinear machine learning model to evaluate their ability to detect informative features. The machine learning models provided flexible variable selection and outperformed linear models in capturing nonlinear dependencies. Model interpretation using SHapley Additive exPlanations (SHAP) indicated that the isoflavone derivative, daidzin, and the drought-tolerant Candidatus Nitrosocosmicus , were major contributors to phenotypic variation, specifically under drought stress. SHAP-based interaction networks indicated cross-omics links, including connections between daidzin, gamma-aminobutyric acid (GABA), and Paenibacillus . Conclusion The proposed interpretable machine learning approach for plant phenotype prediction identified multi-omics biomarkers and interactions, providing insights into plant adaptation to drought stress through environment-dependent rhizosphere networks and symbiotic associations.
Full text 1,752 characters · extracted from oa-doi-fallback · 3 sections · click to expand

Abstract

Background Plant-microbe interactions in the rhizosphere are central to plant growth, nutrient acquisition, and stress resilience. Although multi-omics approaches enable comprehensive profiling of different biological layers, integrating these data to understand the mechanisms underlying plant-microbe symbiosis, particularly under drought stress, remains a challenge.

Results

Genomic, metabolomic, and microbiome data from 198 soybean accessions grown under both control and drought conditions were integrated to identify environment-specific predictive features of the plant phenotypes. We compared best linear unbiased prediction (BLUP), genome-wide association study (GWAS), and a nonlinear machine learning model to evaluate their ability to detect informative features. The machine learning models provided flexible variable selection and outperformed linear models in capturing nonlinear dependencies. Model interpretation using SHapley Additive exPlanations (SHAP) indicated that the isoflavone derivative, daidzin, and the drought-tolerant Candidatus Nitrosocosmicus, were major contributors to phenotypic variation, specifically under drought stress. SHAP-based interaction networks indicated cross-omics links, including connections between daidzin, gamma-aminobutyric acid (GABA), and Paenibacillus.

Conclusion

The proposed interpretable machine learning approach for plant phenotype prediction identified multi-omics biomarkers and interactions, providing insights into plant adaptation to drought stress through environment-dependent rhizosphere networks and symbiotic associations. Competing Interest Statement The authors have declared no competing interest. Footnotes Evaluation methods, writing, authors have been corrected.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00