From Literature to Practice: Application of Large Language Models for Farm Management Insight

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Sustainable food production and security depend on increasing agricultural productivity within existing arable land. This necessitates the effective translation of complex agronomic research into actionable, field-specific crop management recommendations. Despite substantial advances in agricultural research, a persistent knowledge-practice gap continues to impede the widespread adoption of evidence-based management practices. We evaluate whether large language models (LLMs) can bridge this gap by generating crop management recommendations from scientific literature. Using US soybean production as a case study, we developed a semi-automated, human-in-the-loop pipeline adhering to systematic review protocols. Our pipeline demonstrated high accuracy for literature screening, outperforming standalone models. However, when generating a general soybean management plan, expert evaluations rated two commercial LLMs’ output more favorably than the plan from our system. This work highlights the need to develop systems that address user trust and provide tailored, field-specific advice that is both trustworthy and practically useful for farming communities.
Full text 66,593 characters · extracted from preprint-html · click to expand
From Literature to Practice: Application of Large Language Models for Farm Management Insight | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article From Literature to Practice: Application of Large Language Models for Farm Management Insight Spyridon Mourtzinis, Tatiane Severo Silva, Jason Chor Ming Lo, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7424300/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted 10 You are reading this latest preprint version Abstract Sustainable food production and security depend on increasing agricultural productivity within existing arable land. This necessitates the effective translation of complex agronomic research into actionable, field-specific crop management recommendations. Despite substantial advances in agricultural research, a persistent knowledge-practice gap continues to impede the widespread adoption of evidence-based management practices. We evaluate whether large language models (LLMs) can bridge this gap by generating crop management recommendations from scientific literature. Using US soybean production as a case study, we developed a semi-automated, human-in-the-loop pipeline adhering to systematic review protocols. Our pipeline demonstrated high accuracy for literature screening, outperforming standalone models. However, when generating a general soybean management plan, expert evaluations rated two commercial LLMs’ output more favorably than the plan from our system. This work highlights the need to develop systems that address user trust and provide tailored, field-specific advice that is both trustworthy and practically useful for farming communities. Earth and environmental sciences/Environmental sciences Biological sciences/Plant sciences Figures Figure 1 Figure 2 Figure 3 1. Introduction Global population growth, along with rising incomes in developing countries, is expected to exert significant pressure on food production systems worldwide [ 1 ]. Achieving sustainable increases in crop yields will require intensifying production on existing cropland while minimizing ecosystem degradation and greenhouse gas emissions [ 2 ]. Agricultural research plays a crucial role in guiding farmers toward sustainable, productivity-enhancing practices. However, well-documented challenges—such as high environmental variability, fragmented datasets, complex scientific jargon, and limitations in extrapolating findings—hinder the translation of research into actionable recommendations [ 3 ]. Furthermore, the rapid growth in agricultural research output coupled with reductions in Extension and outreach staff exacerbates the challenge of distilling practical insights to support on-farm decision-making. Systematic reviews offer a transparent and rigorous methodology for synthesizing evidence to answer a focused research question [ 4 ]. The process involves defining research questions, developing a search strategy, screening studies against predefined criteria, and extracting and synthesizing data [ 5 ]. The outcomes can help bridge the gap between fragmented information and comprehensive, actionable insights. However, producing such reviews remains a time- and resource-intensive process, leaving a persistent gap between published scientific findings and the practical advice reaching farmers. Recently, the field of artificial intelligence has witnessed rapid advancements in the development of large language models (LLMs) [ 6 , 7 , 8 ]. These models have demonstrated capabilities in answering agriculture-related questions [ 9 ], generating pest management recommendations [ 10 ], and reviewing scientific publications [ 11 ]. These advancements suggest that LLMs could automate time-consuming parts of the knowledge synthesis process. However, current LLM-based tools face challenges, including a lack of control over the knowledge extraction process and the risk of "hallucinations" - where models generate inaccurate or false information [ 12 ]. Such fabrications can negatively affect downstream decision-making, challenging user trust and the adoption of generative AI systems. Here, we explore the use of LLMs to support soybean management planning in the US and evaluate the quality of the generated recommendations through expert review. We focus on US soybean ( Glycine max (L) Merr.) production, which accounts for approximately 30% of global production [ 13 ], as proof of concept. Unlike traditional reviews, we leverage LLMs in a human-in-the-loop workflow to enable a scalable synthesis of agricultural knowledge. Our approach focuses on extracting actionable insights that directly inform farm management decisions, laying the groundwork for efficient, continuously updated ("living") literature reviews in agriculture. 2. Methods We conducted a systematic literature search following established meta-analysis protocols to identify studies relevant to soybean management practices in the US. The workflow we followed is shown in Fig. 1 . The search strategy employed the PICO framework (Population, Intervention, Comparator, Outcome), using three key components: Population (soybeans), Intervention (management practices), and Outcome (yield responses). The search was executed in the Web of Science Core Collection database using the field tag TS (Title, Abstract, Keywords) with publication years restricted to 2015–2025 to capture contemporary agricultural management practices. Search terms were iteratively refined through testing various keyword combinations against known relevant articles to optimize retrieval accuracy (Table S1). Database filters were applied to include only English-language publications and studies conducted within the US. The initial search yielded 990 articles, which were exported in Excel format and accessed via DOI links. Manual screening excluded 92 articles that were not conducted in the US or did not focus on soybean as the primary crop of interest, resulting in a final dataset of 898 articles. All retrieved PDFs were converted to Markdown format using the PyMuPDF Python package to facilitate subsequent automated processing. A multi-model large language model (LLM) screening approach was implemented to evaluate full-text articles against predefined inclusion criteria. Four distinct LLMs (OpenAI-gpt-4.1-mini-2025-04-14, Gemini-2.5-flash-preview-04-17, Deepseek-r1-70b, Llama3.3-70b) independently assessed each study for the following criteria: 1) Geographic scope: Conducted within the US, 2) Crop focus: soybean, 3) Data availability: Includes quantitative yield measurements, 4 ) Experimental setting: Field-based studies (studies performed solely in greenhouse were excluded), 5) Study design: Primary research (meta-analyses and reviews excluded), and 6) Research context: Field experimental conditions. These models were selected through small-scale testing that balanced performance with available computational resources. For each criterion, LLM consensus was required for study inclusion. In cases where LLM assessments resulted in tied decisions affecting downstream inclusion, review by our team (hereafter called “experts”) was conducted to resolve discrepancies and ensure accurate study selection. Additionally, a random sample comprising 20% of the screened articles was independently evaluated by human experts to establish ground truth data. This validation dataset was used to assess: 1) individual LLM performance accuracy, and 2) overall human-in-the-loop pipeline effectiveness. Experts developed ten key research questions addressing critical farm management decisions (Table S2). Given the complexity of these questions, each was systematically decomposed into specific sub-questions using in-context learning with LLMs, guided by two human-provided examples of question decomposition. For example, the question "How do different seed treatments (insecticide or fungicide) impact soybean yield when planted before May 1 compared to after May 1?" was decomposed into: Were insecticide or fungicide seed treatments evaluated in the study? Did the study compare planting timing before versus after May 1? Was soybean yield measured as a primary outcome variable? For each question-paper combination, multiple LLMs independently extracted relevant information, including: 1) determination of study relevance to the specific question, 2) supporting textual evidence from the publication, and 3) synthesized answers based on reported findings. The prompts and output format is shown in tables S4 and S5. A consensus approach was employed whereby papers were considered relevant to a specific question only when at least two LLMs independently identified relevant data. Due to performance issues—specifically, a higher frequency of incorrect or irrelevant answers—we limited this step to the two top-performing models (OpenAI and Gemini). For papers meeting this threshold, extracted answers were processed through a final inconsistency detection algorithm to identify and resolve conflicting interpretations. Human expert review was conducted for all cases where inconsistencies were detected between LLM extractions, ensuring accuracy and reliability of the final synthesized recommendations. Finally, we evaluated the ability of our system and four other AI platforms on June 3, 2025 (OpenAI Deep Research with 4o, Gemini Deep Research with 2.5 Pro, Perplexity Research, Crop Wizard v1.5) to develop a general US-wide soybean farm management plan. The first three are agentic AI systems designed for deep research tasks, capable of autonomously planning, retrieving, and synthesizing information to answer complex questions. In contrast, Crop Wizard is a retrieval-augmented generation (RAG) system tailored specifically for agricultural use cases. Each system was given the prompt: “Create a soybean farm management plan tailored to U.S. conditions.”. The resulting plans were evaluated through a survey completed by 44 independent soybean experts from the U.S., including university faculty, state soybean Extension specialists, and representatives from state or national soybean boards. Experts rated the correctness, completeness, practicality, and relevance of each plan and provided an overall rating. The survey also collected data on expert sentiment regarding the use of AI in agriculture. 3. Results The four LLMs demonstrated varied performance in screening research articles against inclusion criteria (Table 1). The commercial models (OpenAI and Gemini) generally outperformed the open-source models (Llama and DeepSeek), achieving F1 scores ranging from 0.81 to 0.99 across the criteria. The open-source models had a wider performance range, with F1 scores between 0.43 and 0.92. The human-in-the-loop approach, which used expert review to resolve ties, consistently achieved the highest F1 scores, reaching near-perfect or perfect scores (0.98 to 1.00) for all criteria. The most challenging criterion for all individual LLMs was identifying studies that were not conducted in a greenhouse, where F1 scores were as low as 0.43 (Llama) and 0.48 (DeepSeek). The human-in-the-loop method successfully resolved these ambiguities, achieving an F1 score of 1.00 for this criterion and an overall combined screening F1 score of 0.96, surpassing the best individual model (OpenAI, 0.90). Table 1. Comparison of F1 Scores for Human-in-the-Loop and LLM approaches in screening research articles. The F1 score (ranging from 0 to 1) measures a model's accuracy by calculating the harmonic mean of its precision and recall. Precision is the proportion of positive identifications that were actually correct, while recall is the proportion of actual positives that were correctly identified. A higher F1 score indicates better classification performance. Screening Criteria Llama DeepSeek OpenAI Gemini Human-in-the-loop Is primary study 0.82 0.82 0.98 0.99 0.99 Soybean-focused 0.92 0.91 0.97 0.97 0.99 Yield data collected 0.82 0.80 0.92 0.94 0.98 Conducted in the US 0.84 0.92 0.99 0.99 1.00 Is field study 0.92 0.92 0.94 0.97 0.98 Not conducted in greenhouse 0.43 0.48 0.81 0.81 1.00 Combined Screening Results 0.71 0.70 0.90 0.89 0.96 Note. The human-in-the-loop method uses a majority vote from the four LLMs, with a human expert serving as the tiebreaker. The soybean management plans generated by the five AI systems were rated differently by the 44 domain experts (Figure 2). Based on the expert evaluations, Gemini's plan received the highest overall ratings, with approximately 79% of experts rating it as 'Good' or 'Excellent'. The OpenAI plan followed with about 69% positive ratings, and our system's plan was rated 'Good' or 'Excellent' by approximately 68% of experts. The plan from Perplexity received the lowest overall positive rating at 57%, while Crop Wizard received an overall rating of 66%. Specifically, for 'Correctness', Gemini's plan was again the top performer, with 88% of experts giving it a 'Good' or 'Excellent' rating. Our system's plan was also rated highly for correctness, receiving a combined 'Good' or 'Excellent' rating from 77% of experts, just ahead of OpenAI's plan (76%). For 'Completeness', Gemini's plan was rated highest, with about 82% of experts rating it 'Good' or 'Excellent'. Our system's plan received a lower positive rating for completeness (55%). Regarding 'Practicality', Gemini's plan was also perceived most favorably, with over 74% of experts giving it 'Good' or 'Excellent' ratings. In contrast, our system's plan received the highest percentage of 'Poor' ratings for practicality, at nearly 19%. No system was entirely free of negative ratings, as every system received at least one 'Poor' or 'Very poor' rating on some metric. These negative evaluations can likely be attributed to LLM errors or the omission of key aspects of soybean management. It is critical to note that the management plans were generated from a prompt requesting general recommendations applicable across the US, rather than for specific agroecological regions. Consequently, the omission of region-specific details, which are critically important to agricultural experts, likely contributed to the negative ratings. For example, an optimal planting date or pest management strategy for soybeans in the north-central US would differ significantly from best practices in the eastern US. Therefore, a generalized recommendation, while factually correct on a broad scale, could be rated impractical or incomplete by an expert evaluating it for their specific locale. The survey of agricultural experts revealed a generally positive but nuanced sentiment toward the application of AI in their field (Figure 3). There was strong agreement on the potential benefits for increased efficiency, with approximately 77% of respondents selecting 'Strongly agree' or 'Somewhat agree'. Sentiment regarding increased usefulness was also positive, with about 59% of experts agreeing. However, sentiment regarding trust in AI was more divided: only about 41% of experts agreed that AI is trustworthy, while a similar proportion (39%) remained neutral ('Neither agree nor disagree'), and 20% expressed some level of distrust. Opinions on data sharing for AI development were also mixed. While a majority (about 57%) agreed to share data, approximately 25% were neutral, and nearly 18% were hesitant, expressing disagreement. This suggests that while experts recognize the potential utility of AI, building trust and addressing data-sharing concerns are critical for its broader adoption. 4. Discussion This study demonstrates the potential and current limitations of using LLMs to bridge the gap between scientific research and practical farm management. Our results from the initial literature screening process show that a human-in-the-loop approach is currently essential for achieving high accuracy, significantly outperforming the best standalone model. This greater performance was most evident in nuanced tasks, such as differentiating between field and greenhouse studies. This highlights that while LLMs can greatly accelerate the initial stages of a systematic review [ 14 ], human expertise remains indispensable for ensuring the quality and accuracy of the selected evidence base [ 15 ]. In the generation of farm management plans, the outcomes were more complex. The plan generated by the standalone Gemini model received the most favorable ratings from experts, particularly for correctness and completeness. In contrast, our system, despite being built on a more rigorously screened dataset, received lower ratings for completeness and practicality. This finding suggests that the data corpus we synthesized was not large and diverse enough to capture all specific aspects of soybean management ( e.g. crop marketing). Additionally, the final synthesis and presentation of information are as important as the quality of the underlying data. Generating a single plan for the entire US likely omitted the crucial region-specific details that agricultural experts rely on for making practical, on-the-ground decisions. This underscores a key challenge for developing effective agricultural AI: the need to move beyond generalized knowledge to location-specific [ 3 ], actionable intelligence with some level of human oversight. The expert sentiment analysis further contextualizes these findings. While a strong majority of experts recognized the potential for AI to improve efficiency (77% agreement) and usefulness (58% agreement), there were significant reservations regarding trust and data sharing which aligns with farmers reluctance to share data [ 16 ]. Only 40% of experts expressed trust in AI, and nearly 19% were hesitant to share data. This "trust deficit" is a major barrier to adoption. It indicates that even a technically perfect AI system will fail to have an impact if the intended users do not have confidence in its recommendations or the governance of its data. 5. Conclusions Our findings show that LLMs, when guided by human experts, can serve as powerful tools to accelerate the synthesis of complex agricultural research. However, the generation of valuable farm management advice requires a deep understanding of context and practical application. The failure of a generalized plan to satisfy experts highlights that the future of agricultural AI lies in its ability to provide tailored, region or in most cases field-specific recommendations. The success of these technologies will ultimately depend on building trust within the agricultural community. Future work should focus not only on improving the technical accuracy and specificity of AI models but also on creating transparent, secure, and collaborative systems that empower, rather than replace, human expertise. This study represents a crucial step in navigating the path toward an AI-assisted future for agriculture that balances technological potential with practical needs and inherent skepticism of its intended users. Declarations Author Contribution D.S. and S.C. contributed to Conceptualization, Project Administration, and Writing – Review & Editing. S.M., J.L., and T.S. were responsible for Data Curation. J.L. performed the Formal Analysis. S.M. prepared the Writing – Original Draft, and J.L., T.S., D.S., and S.C. contributed to Writing – Review & Editing. Data availability The data can be made available upon request. Contact corresponding author with any queries. Code availability All LLM prompts used are provided in supplementary material Funding Declaration The authors declare that no funding was received for this study. References Searchinger, T., Waite, R., Hanson, C., Ranganathan, J., Dumas, P., 2019. Creating a Sustainable Food Future: A Menu of Solutions to Feed Nearly 10 Billion People by 2050. World Resources Institute, Washington, D.C. Foley, J.A., Ramankutty, N., Brauman, K.A., Cassidy, E.S., Gerber, J.S., Johnston, M., Mueller, N.D., O’Connell, C., Ray, D.K., West, P.C., Balzer, C., Bennett, E.M., Carpenter, S.R., Hill, J., Monfreda, C., Polasky, S., Rockström, J., Sheehan, J., Siebert, D., Tilman, D., Zaks, D.P.M., 2011. Solutions for a cultivated planet. Nature 478, 337–342. Mourtzinis, S., Esker, P.D., Specht, J.E., Conley, S.P., 2021. Advancing agricultural research using machine learning algorithms. Sci. Rep. 11, 17879. van der Knaap, L.M., 2008. Combining Campbell standard and the realist evaluation approach: the best of two worlds? Am. J. Eval. 29, 48–57. Mallett, R., Hagen-Zanker, J., Slater, R., Duvendack, M., 2012. The benefits and challenges of using systematic reviews in international development research. J. Dev. Eff. 4, 445–455. OpenAI, 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774. AI@Meta, 2024. Llama 3 Model Card. Meta AI. Anthropic, 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku. Anthropic. Silva, B., Nunes, L., Estevão, R., Aski, V., Chandra, R., 2023. GPT-4 as an agronomist assistant? Answering agriculture exams using large language models. arXiv preprint arXiv:2310.06225. Yang, S., Yuan, Z., Li, S., Peng, R., Liu, K., Yang, P., 2024. GPT-4 as evaluator: evaluating large language models on pest management in agriculture. arXiv preprint arXiv:2403.11858. Zhu, M., Weng, Y., Yang, L., Zhang, Y., 2024. DeepReview: improving LLM-based paper review with human-like deep thinking process. arXiv preprint arXiv:2403.08569. Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., Ruetsch-Chelli, C., 2024. Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: comparative analysis. J. Med. Internet Res. 26, e53164. Food and Agriculture Organization of the United Nations, 2023. FAOSTAT database: crops and livestock products. Available at: https://www.fao.org/faostat/en/#data/QCL (Accessed 8 August 2025). van Dijk, S.H.B., Brusse-Keizer, M.G.J., Bucsán, C.C., van der Palen, J., Doggen, C.J.M., Lenferink, A., 2023. Artificial intelligence in systematic reviews: promising when appropriately used. BMJ Open 13, e072254. Blaizot, A., Veettil, S.K., Saidoung, P., Moreno-Garcia, C.F., Wiratunga, N., Aceves-Martins, M., Lai, N.M., Chaiyakunapruk, N., 2022. Using artificial intelligence methods for systematic review in health sciences: a systematic review. Res. Synth. Methods 13, 289–303. Wiseman, L., Sanderson, J., Zhang, A., Jakku, E., 2019. Farmers and their data: an examination of farmers' reluctance to share their data for the purposes of digital agriculture. NJAS Wagening. J. Life Sci. 90–91, 100301. Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterial.docx Cite Share Download PDF Status: Published Journal Publication published 02 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 10 Oct, 2025 Reviews received at journal 29 Sep, 2025 Reviews received at journal 23 Sep, 2025 Reviewers agreed at journal 08 Sep, 2025 Reviewers agreed at journal 08 Sep, 2025 Reviewers invited by journal 03 Sep, 2025 Editor assigned by journal 03 Sep, 2025 Editor invited by journal 29 Aug, 2025 Submission checks completed at journal 26 Aug, 2025 First submitted to journal 26 Aug, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7424300","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":511787691,"identity":"fd5927aa-a308-42c4-97fc-ecea77e9af3c","order_by":0,"name":"Spyridon Mourtzinis","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA30lEQVRIiWNgGAWjYPACCR429gYgbWBBlHLGBoYEGzk+ngMgLRJEa0kzlpNIAFtHWD1///LnD37+OJzYJvn86oYfBRIM/O3dCXi1SNx4Y9jYkwDUIp1TdrMH6DCJM2c34LfmxhnGBh6IlrQbPEAtBhK5+LXI3zj+sPEPSIvkmbSbf4jRYnC+wbCZB+h9Ngn2Y7eJssXwBo/hbJk0Gzk2nhy22zIGEjwE/SJ3/viDj29sJHjk248/u/nmj40cf3svAe9DogMEeAzAJH7lIMB/AMZif0BY9SgYBaNgFIxIAABwUkrTHyXC9wAAAABJRU5ErkJggg==","orcid":"","institution":"University of Wisconsin-Madison","correspondingAuthor":true,"prefix":"","firstName":"Spyridon","middleName":"","lastName":"Mourtzinis","suffix":""},{"id":511787692,"identity":"dcb74eeb-e3f1-417c-8934-09badfae369a","order_by":1,"name":"Tatiane Severo Silva","email":"","orcid":"","institution":"University of Wisconsin-Madison","correspondingAuthor":false,"prefix":"","firstName":"Tatiane","middleName":"Severo","lastName":"Silva","suffix":""},{"id":511787693,"identity":"d34d37c2-9e75-4150-ac21-028956143717","order_by":2,"name":"Jason Chor Ming Lo","email":"","orcid":"","institution":"University of Wisconsin-Madison","correspondingAuthor":false,"prefix":"","firstName":"Jason","middleName":"Chor Ming","lastName":"Lo","suffix":""},{"id":511787694,"identity":"0f0a86f2-01af-499e-bdac-b2385fa79329","order_by":3,"name":"Damon L. Smith","email":"","orcid":"","institution":"University of Wisconsin-Madison","correspondingAuthor":false,"prefix":"","firstName":"Damon","middleName":"L.","lastName":"Smith","suffix":""},{"id":511787695,"identity":"eda42fcd-61ef-45aa-9b7d-35c24bcd50bf","order_by":4,"name":"Shawn P. Conley","email":"","orcid":"","institution":"University of Wisconsin-Madison","correspondingAuthor":false,"prefix":"","firstName":"Shawn","middleName":"P.","lastName":"Conley","suffix":""}],"badges":[],"createdAt":"2025-08-21 09:08:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7424300/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7424300/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-30991-6","type":"published","date":"2025-12-02T15:57:27+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":90929237,"identity":"6e09b69a-45d7-49ad-b91f-bc9990a7b3d6","added_by":"auto","created_at":"2025-09-09 16:04:30","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":529422,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow description from screening, distilling and summarizing data from research studies. The screening process selects PDFs that meet the inclusion criteria. The distilling process extracts relevant knowledge for each question-PDF pair. The summarizing step generates an expert-guided, fully traceable knowledge summary for soybean farm management planning. Large language model (LLM) tasks are shown in yellow, and human tasks are shown in blue.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7424300/v1/9220e1a8147317584914161f.jpeg"},{"id":90930642,"identity":"4830bfdc-73a2-4a39-bb90-a70f8693e3a6","added_by":"auto","created_at":"2025-09-09 16:12:30","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":714685,"visible":true,"origin":"","legend":"\u003cp\u003eEvaluation of soybean management plans tailored to US conditions generated by OpenAI, Gemini, Perplexity, Crop Wizard and our system. Management plans were evaluated by 44 soybean research experts across the north central US via survey questionnaire.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7424300/v1/dea001dc255dfc9958970183.jpeg"},{"id":90929235,"identity":"0e372b64-bbbc-466d-b5ee-8ed58c66e733","added_by":"auto","created_at":"2025-09-09 16:04:30","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":111171,"visible":true,"origin":"","legend":"\u003cp\u003eSentiment of agricultural experts towards the application of artificial intelligence in farming, based on survey responses.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7424300/v1/6b9ec860be1ecd80349f6147.png"},{"id":97724238,"identity":"3eac4d4a-b521-458d-af02-8776561b6d5d","added_by":"auto","created_at":"2025-12-08 16:12:16","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1729783,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7424300/v1/a0b54873-62ca-4c3d-8292-bb5899c34d2f.pdf"},{"id":90929236,"identity":"c50ad471-6ba5-4ae6-8a4d-96d225b9758b","added_by":"auto","created_at":"2025-09-09 16:04:30","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":20111,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-7424300/v1/17dacb376481269e3a5aaecc.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"From Literature to Practice: Application of Large Language Models for Farm Management Insight","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eGlobal population growth, along with rising incomes in developing countries, is expected to exert significant pressure on food production systems worldwide [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Achieving sustainable increases in crop yields will require intensifying production on existing cropland while minimizing ecosystem degradation and greenhouse gas emissions [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Agricultural research plays a crucial role in guiding farmers toward sustainable, productivity-enhancing practices. However, well-documented challenges\u0026mdash;such as high environmental variability, fragmented datasets, complex scientific jargon, and limitations in extrapolating findings\u0026mdash;hinder the translation of research into actionable recommendations [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Furthermore, the rapid growth in agricultural research output coupled with reductions in Extension and outreach staff exacerbates the challenge of distilling practical insights to support on-farm decision-making.\u003c/p\u003e\u003cp\u003eSystematic reviews offer a transparent and rigorous methodology for synthesizing evidence to answer a focused research question [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The process involves defining research questions, developing a search strategy, screening studies against predefined criteria, and extracting and synthesizing data [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. The outcomes can help bridge the gap between fragmented information and comprehensive, actionable insights. However, producing such reviews remains a time- and resource-intensive process, leaving a persistent gap between published scientific findings and the practical advice reaching farmers.\u003c/p\u003e\u003cp\u003eRecently, the field of artificial intelligence has witnessed rapid advancements in the development of large language models (LLMs) [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. These models have demonstrated capabilities in answering agriculture-related questions [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], generating pest management recommendations [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], and reviewing scientific publications [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. These advancements suggest that LLMs could automate time-consuming parts of the knowledge synthesis process. However, current LLM-based tools face challenges, including a lack of control over the knowledge extraction process and the risk of \"hallucinations\" - where models generate inaccurate or false information [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Such fabrications can negatively affect downstream decision-making, challenging user trust and the adoption of generative AI systems.\u003c/p\u003e\u003cp\u003eHere, we explore the use of LLMs to support soybean management planning in the US and evaluate the quality of the generated recommendations through expert review. We focus on US soybean (\u003cem\u003eGlycine max\u003c/em\u003e (L) Merr.) production, which accounts for approximately 30% of global production [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], as proof of concept. Unlike traditional reviews, we leverage LLMs in a human-in-the-loop workflow to enable a scalable synthesis of agricultural knowledge. Our approach focuses on extracting actionable insights that directly inform farm management decisions, laying the groundwork for efficient, continuously updated (\"living\") literature reviews in agriculture.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cp\u003eWe conducted a systematic literature search following established meta-analysis protocols to identify studies relevant to soybean management practices in the US. The workflow we followed is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The search strategy employed the PICO framework (Population, Intervention, Comparator, Outcome), using three key components: Population (soybeans), Intervention (management practices), and Outcome (yield responses).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe search was executed in the Web of Science Core Collection database using the field tag TS (Title, Abstract, Keywords) with publication years restricted to 2015\u0026ndash;2025 to capture contemporary agricultural management practices. Search terms were iteratively refined through testing various keyword combinations against known relevant articles to optimize retrieval accuracy (Table S1). Database filters were applied to include only English-language publications and studies conducted within the US. The initial search yielded 990 articles, which were exported in Excel format and accessed via DOI links. Manual screening excluded 92 articles that were not conducted in the US or did not focus on soybean as the primary crop of interest, resulting in a final dataset of 898 articles. All retrieved PDFs were converted to Markdown format using the PyMuPDF Python package to facilitate subsequent automated processing.\u003c/p\u003e\u003cp\u003eA multi-model large language model (LLM) screening approach was implemented to evaluate full-text articles against predefined inclusion criteria. Four distinct LLMs (OpenAI-gpt-4.1-mini-2025-04-14, Gemini-2.5-flash-preview-04-17, Deepseek-r1-70b, Llama3.3-70b) independently assessed each study for the following criteria: 1) Geographic scope: Conducted within the US, 2) Crop focus: soybean, 3) Data availability: Includes quantitative yield measurements, 4 ) Experimental setting: Field-based studies (studies performed solely in greenhouse were excluded), 5) Study design: Primary research (meta-analyses and reviews excluded), and 6) Research context: Field experimental conditions. These models were selected through small-scale testing that balanced performance with available computational resources.\u003c/p\u003e\u003cp\u003eFor each criterion, LLM consensus was required for study inclusion. In cases where LLM assessments resulted in tied decisions affecting downstream inclusion, review by our team (hereafter called \u0026ldquo;experts\u0026rdquo;) was conducted to resolve discrepancies and ensure accurate study selection. Additionally, a random sample comprising 20% of the screened articles was independently evaluated by human experts to establish ground truth data. This validation dataset was used to assess: 1) individual LLM performance accuracy, and 2) overall human-in-the-loop pipeline effectiveness.\u003c/p\u003e\u003cp\u003eExperts developed ten key research questions addressing critical farm management decisions (Table S2). Given the complexity of these questions, each was systematically decomposed into specific sub-questions using in-context learning with LLMs, guided by two human-provided examples of question decomposition. For example, the question \"How do different seed treatments (insecticide or fungicide) impact soybean yield when planted before May 1 compared to after May 1?\" was decomposed into: Were insecticide or fungicide seed treatments evaluated in the study? Did the study compare planting timing before versus after May 1? Was soybean yield measured as a primary outcome variable? For each question-paper combination, multiple LLMs independently extracted relevant information, including: 1) determination of study relevance to the specific question, 2) supporting textual evidence from the publication, and 3) synthesized answers based on reported findings. The prompts and output format is shown in tables S4 and S5.\u003c/p\u003e\u003cp\u003eA consensus approach was employed whereby papers were considered relevant to a specific question only when at least two LLMs independently identified relevant data. Due to performance issues\u0026mdash;specifically, a higher frequency of incorrect or irrelevant answers\u0026mdash;we limited this step to the two top-performing models (OpenAI and Gemini). For papers meeting this threshold, extracted answers were processed through a final inconsistency detection algorithm to identify and resolve conflicting interpretations. Human expert review was conducted for all cases where inconsistencies were detected between LLM extractions, ensuring accuracy and reliability of the final synthesized recommendations.\u003c/p\u003e\u003cp\u003eFinally, we evaluated the ability of our system and four other AI platforms on June 3, 2025 (OpenAI Deep Research with 4o, Gemini Deep Research with 2.5 Pro, Perplexity Research, Crop Wizard v1.5) to develop a general US-wide soybean farm management plan. The first three are agentic AI systems designed for deep research tasks, capable of autonomously planning, retrieving, and synthesizing information to answer complex questions. In contrast, Crop Wizard is a retrieval-augmented generation (RAG) system tailored specifically for agricultural use cases. Each system was given the prompt: \u0026ldquo;Create a soybean farm management plan tailored to U.S. conditions.\u0026rdquo;. The resulting plans were evaluated through a survey completed by 44 independent soybean experts from the U.S., including university faculty, state soybean Extension specialists, and representatives from state or national soybean boards. Experts rated the correctness, completeness, practicality, and relevance of each plan and provided an overall rating. The survey also collected data on expert sentiment regarding the use of AI in agriculture.\u003c/p\u003e"},{"header":"3. Results","content":"\u003cp\u003eThe four LLMs demonstrated varied performance in screening research articles against inclusion criteria (Table 1). The commercial models (OpenAI and Gemini) generally outperformed the open-source models (Llama and DeepSeek), achieving F1 scores ranging from 0.81 to 0.99 across the criteria. The open-source models had a wider performance range, with F1 scores between 0.43 and 0.92. The human-in-the-loop approach, which used expert review to resolve ties, consistently achieved the highest F1 scores, reaching near-perfect or perfect scores (0.98 to 1.00) for all criteria. The most challenging criterion for all individual LLMs was identifying studies that were not conducted in a greenhouse, where F1 scores were as low as 0.43 (Llama) and 0.48 (DeepSeek). The human-in-the-loop method successfully resolved these ambiguities, achieving an F1 score of 1.00 for this criterion and an overall combined screening F1 score of 0.96, surpassing the best individual model (OpenAI, 0.90).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1.\u0026nbsp;\u003c/strong\u003eComparison of F1 Scores for Human-in-the-Loop and LLM approaches in screening research articles. The F1 score (ranging from 0 to 1) measures a model\u0026apos;s accuracy by calculating the harmonic mean of its precision and recall. Precision is the proportion of positive identifications that were actually correct, while recall is the proportion of actual positives that were correctly identified. A higher F1 score indicates better classification performance.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eScreening Criteria\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003eLlama\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003eDeepSeek\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003eOpenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003eGemini\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003eHuman-in-the-loop\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eIs primary study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.98\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e0.99\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eSoybean-focused\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e0.99\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eYield data collected\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e0.98\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eConducted in the US\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eIs field study\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e0.98\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eNot conducted in greenhouse\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.43\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.48\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"bottom\" style=\"width: 35%;\"\u003e\n \u003cp\u003eCombined Screening Results\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 9.16667%;\"\u003e\n \u003cp\u003e0.71\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 13.1667%;\"\u003e\n \u003cp\u003e0.70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 8.66667%;\"\u003e\n \u003cp\u003e0.90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 10.3333%;\"\u003e\n \u003cp\u003e0.89\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"bottom\" style=\"width: 23.6667%;\"\u003e\n \u003cp\u003e0.96\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eNote. The human-in-the-loop method uses a majority vote from the four LLMs, with a human expert serving as the tiebreaker.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe soybean management plans generated by the five AI systems were rated differently by the 44 domain experts (Figure 2). Based on the expert evaluations, Gemini\u0026apos;s plan received the highest overall ratings, with approximately 79% of experts rating it as \u0026apos;Good\u0026apos; or \u0026apos;Excellent\u0026apos;. The OpenAI plan followed with about 69% positive ratings, and our system\u0026apos;s plan was rated \u0026apos;Good\u0026apos; or \u0026apos;Excellent\u0026apos; by approximately 68% of experts. The plan from Perplexity received the lowest overall positive rating at 57%, while Crop Wizard received an overall rating of 66%.\u003c/p\u003e\n\u003cp\u003eSpecifically, for \u0026apos;Correctness\u0026apos;, Gemini\u0026apos;s plan was again the top performer, with 88% of experts giving it a \u0026apos;Good\u0026apos; or \u0026apos;Excellent\u0026apos; rating. Our system\u0026apos;s plan was also rated highly for correctness, receiving a combined \u0026apos;Good\u0026apos; or \u0026apos;Excellent\u0026apos; rating from 77% of experts, just ahead of OpenAI\u0026apos;s plan (76%). For \u0026apos;Completeness\u0026apos;, Gemini\u0026apos;s plan was rated highest, with about 82% of experts rating it \u0026apos;Good\u0026apos; or \u0026apos;Excellent\u0026apos;. Our system\u0026apos;s plan received a lower positive rating for completeness (55%). Regarding \u0026apos;Practicality\u0026apos;, Gemini\u0026apos;s plan was also perceived most favorably, with over 74% of experts giving it \u0026apos;Good\u0026apos; or \u0026apos;Excellent\u0026apos; ratings. In contrast, our system\u0026apos;s plan received the highest percentage of \u0026apos;Poor\u0026apos; ratings for practicality, at nearly 19%.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eNo system was entirely free of negative ratings, as every system received at least one \u0026apos;Poor\u0026apos; or \u0026apos;Very poor\u0026apos; rating on some metric. These negative evaluations can likely be attributed to LLM errors or the omission of key aspects of soybean management. It is critical to note that the management plans were generated from a prompt requesting general recommendations applicable across the US, rather than for specific agroecological regions. Consequently, the omission of region-specific details, which are critically important to agricultural experts, likely contributed to the negative ratings. For example, an optimal planting date or pest management strategy for soybeans in the north-central US would differ significantly from best practices in the eastern US. Therefore, a generalized recommendation, while factually correct on a broad scale, could be rated impractical or incomplete by an expert evaluating it for their specific locale.\u003c/p\u003e\n\u003cp\u003eThe survey of agricultural experts revealed a generally positive but nuanced sentiment toward the application of AI in their field (Figure 3). There was strong agreement on the potential benefits for increased efficiency, with approximately 77% of respondents selecting \u0026apos;Strongly agree\u0026apos; or \u0026apos;Somewhat agree\u0026apos;. Sentiment regarding increased usefulness was also positive, with about 59% of experts agreeing. However, sentiment regarding trust in AI was more divided: only about 41% of experts agreed that AI is trustworthy, while a similar proportion (39%) remained neutral (\u0026apos;Neither agree nor disagree\u0026apos;), and 20% expressed some level of distrust. Opinions on data sharing for AI development were also mixed. While a majority (about 57%) agreed to share data, approximately 25% were neutral, and nearly 18% were hesitant, expressing disagreement. This suggests that while experts recognize the potential utility of AI, building trust and addressing data-sharing concerns are critical for its broader adoption.\u003c/p\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis study demonstrates the potential and current limitations of using LLMs to bridge the gap between scientific research and practical farm management. Our results from the initial literature screening process show that a human-in-the-loop approach is currently essential for achieving high accuracy, significantly outperforming the best standalone model. This greater performance was most evident in nuanced tasks, such as differentiating between field and greenhouse studies. This highlights that while LLMs can greatly accelerate the initial stages of a systematic review [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], human expertise remains indispensable for ensuring the quality and accuracy of the selected evidence base [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eIn the generation of farm management plans, the outcomes were more complex. The plan generated by the standalone Gemini model received the most favorable ratings from experts, particularly for correctness and completeness. In contrast, our system, despite being built on a more rigorously screened dataset, received lower ratings for completeness and practicality. This finding suggests that the data corpus we synthesized was not large and diverse enough to capture all specific aspects of soybean management (\u003cem\u003ee.g.\u003c/em\u003e crop marketing). Additionally, the final synthesis and presentation of information are as important as the quality of the underlying data. Generating a single plan for the entire US likely omitted the crucial region-specific details that agricultural experts rely on for making practical, on-the-ground decisions. This underscores a key challenge for developing effective agricultural AI: the need to move beyond generalized knowledge to location-specific [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], actionable intelligence with some level of human oversight.\u003c/p\u003e\u003cp\u003eThe expert sentiment analysis further contextualizes these findings. While a strong majority of experts recognized the potential for AI to improve efficiency (77% agreement) and usefulness (58% agreement), there were significant reservations regarding trust and data sharing which aligns with farmers reluctance to share data [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Only 40% of experts expressed trust in AI, and nearly 19% were hesitant to share data. This \"trust deficit\" is a major barrier to adoption. It indicates that even a technically perfect AI system will fail to have an impact if the intended users do not have confidence in its recommendations or the governance of its data.\u003c/p\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eOur findings show that LLMs, when guided by human experts, can serve as powerful tools to accelerate the synthesis of complex agricultural research. However, the generation of valuable farm management advice requires a deep understanding of context and practical application. The failure of a generalized plan to satisfy experts highlights that the future of agricultural AI lies in its ability to provide tailored, region or in most cases field-specific recommendations. The success of these technologies will ultimately depend on building trust within the agricultural community. Future work should focus not only on improving the technical accuracy and specificity of AI models but also on creating transparent, secure, and collaborative systems that empower, rather than replace, human expertise. This study represents a crucial step in navigating the path toward an AI-assisted future for agriculture that balances technological potential with practical needs and inherent skepticism of its intended users.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eD.S. and S.C. contributed to Conceptualization, Project Administration, and Writing \u0026ndash; Review \u0026amp; Editing. S.M., J.L., and T.S. were responsible for Data Curation. J.L. performed the Formal Analysis. S.M. prepared the Writing \u0026ndash; Original Draft, and J.L., T.S., D.S., and S.C. contributed to Writing \u0026ndash; Review \u0026amp; Editing.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data can be made available upon request. Contact corresponding author with any queries.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll LLM prompts used are provided in supplementary material\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;Funding Declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that no funding was received for this study.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSearchinger, T., Waite, R., Hanson, C., Ranganathan, J., Dumas, P., 2019. Creating a Sustainable Food Future: A Menu of Solutions to Feed Nearly 10 Billion People by 2050. World Resources Institute, Washington, D.C.\u003c/li\u003e\n\u003cli\u003eFoley, J.A., Ramankutty, N., Brauman, K.A., Cassidy, E.S., Gerber, J.S., Johnston, M., Mueller, N.D., O\u0026rsquo;Connell, C., Ray, D.K., West, P.C., Balzer, C., Bennett, E.M., Carpenter, S.R., Hill, J., Monfreda, C., Polasky, S., Rockstr\u0026ouml;m, J., Sheehan, J., Siebert, D., Tilman, D., Zaks, D.P.M., 2011. Solutions for a cultivated planet. Nature 478, 337\u0026ndash;342.\u003c/li\u003e\n\u003cli\u003eMourtzinis, S., Esker, P.D., Specht, J.E., Conley, S.P., 2021. Advancing agricultural research using machine learning algorithms. Sci. Rep. 11, 17879.\u003c/li\u003e\n\u003cli\u003evan der Knaap, L.M., 2008. Combining Campbell standard and the realist evaluation approach: the best of two worlds? Am. J. Eval. 29, 48\u0026ndash;57.\u003c/li\u003e\n\u003cli\u003eMallett, R., Hagen-Zanker, J., Slater, R., Duvendack, M., 2012. The benefits and challenges of using systematic reviews in international development research. J. Dev. Eff. 4, 445\u0026ndash;455.\u003c/li\u003e\n\u003cli\u003eOpenAI, 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774.\u003c/li\u003e\n\u003cli\u003eAI@Meta, 2024. Llama 3 Model Card. Meta AI.\u003c/li\u003e\n\u003cli\u003eAnthropic, 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku. Anthropic.\u003c/li\u003e\n\u003cli\u003eSilva, B., Nunes, L., Estev\u0026atilde;o, R., Aski, V., Chandra, R., 2023. GPT-4 as an agronomist assistant? Answering agriculture exams using large language models. arXiv preprint arXiv:2310.06225.\u003c/li\u003e\n\u003cli\u003eYang, S., Yuan, Z., Li, S., Peng, R., Liu, K., Yang, P., 2024. GPT-4 as evaluator: evaluating large language models on pest management in agriculture. arXiv preprint arXiv:2403.11858.\u003c/li\u003e\n\u003cli\u003eZhu, M., Weng, Y., Yang, L., Zhang, Y., 2024. DeepReview: improving LLM-based paper review with human-like deep thinking process. arXiv preprint arXiv:2403.08569.\u003c/li\u003e\n\u003cli\u003eChelli, M., Descamps, J., Lavou\u0026eacute;, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., Ruetsch-Chelli, C., 2024. Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: comparative analysis. J. Med. Internet Res. 26, e53164.\u003c/li\u003e\n\u003cli\u003eFood and Agriculture Organization of the United Nations, 2023. FAOSTAT database: crops and livestock products. Available at: https://www.fao.org/faostat/en/#data/QCL (Accessed 8 August 2025).\u003c/li\u003e\n\u003cli\u003evan Dijk, S.H.B., Brusse-Keizer, M.G.J., Bucs\u0026aacute;n, C.C., van der Palen, J., Doggen, C.J.M., Lenferink, A., 2023. Artificial intelligence in systematic reviews: promising when appropriately used. BMJ Open 13, e072254.\u003c/li\u003e\n\u003cli\u003eBlaizot, A., Veettil, S.K., Saidoung, P., Moreno-Garcia, C.F., Wiratunga, N., Aceves-Martins, M., Lai, N.M., Chaiyakunapruk, N., 2022. Using artificial intelligence methods for systematic review in health sciences: a systematic review. Res. Synth. Methods 13, 289\u0026ndash;303.\u003c/li\u003e\n\u003cli\u003eWiseman, L., Sanderson, J., Zhang, A., Jakku, E., 2019. Farmers and their data: an examination of farmers\u0026apos; reluctance to share their data for the purposes of digital agriculture. NJAS Wagening. J. Life Sci. 90\u0026ndash;91, 100301.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7424300/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7424300/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSustainable food production and security depend on increasing agricultural productivity within existing arable land. This necessitates the effective translation of complex agronomic research into actionable, field-specific crop management recommendations. Despite substantial advances in agricultural research, a persistent knowledge-practice gap continues to impede the widespread adoption of evidence-based management practices. We evaluate whether large language models (LLMs) can bridge this gap by generating crop management recommendations from scientific literature. Using US soybean production as a case study, we developed a semi-automated, human-in-the-loop pipeline adhering to systematic review protocols. Our pipeline demonstrated high accuracy for literature screening, outperforming standalone models. However, when generating a general soybean management plan, expert evaluations rated two commercial LLMs\u0026rsquo; output more favorably than the plan from our system. This work highlights the need to develop systems that address user trust and provide tailored, field-specific advice that is both trustworthy and practically useful for farming communities.\u003c/p\u003e","manuscriptTitle":"From Literature to Practice: Application of Large Language Models for Farm Management Insight","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-09 16:04:26","doi":"10.21203/rs.3.rs-7424300/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-10-10T18:44:08+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-30T02:33:35+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-23T19:00:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"89287793200933035496648421938830685981","date":"2025-09-08T04:55:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"213589480909983526542434402360785536070","date":"2025-09-08T04:09:01+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-09-03T04:07:20+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-09-03T04:06:01+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-08-29T07:02:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-08-26T06:57:30+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-08-26T06:54:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7fc6bcac-d4ac-4637-8626-a5ace7f962bf","owner":[],"postedDate":"September 9th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":54350559,"name":"Earth and environmental sciences/Environmental sciences"},{"id":54350560,"name":"Biological sciences/Plant sciences"}],"tags":[],"updatedAt":"2025-12-08T16:07:58+00:00","versionOfRecord":{"articleIdentity":"rs-7424300","link":"https://doi.org/10.1038/s41598-025-30991-6","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-12-02 15:57:27","publishedOnDateReadable":"December 2nd, 2025"},"versionCreatedAt":"2025-09-09 16:04:26","video":"","vorDoi":"10.1038/s41598-025-30991-6","vorDoiUrl":"https://doi.org/10.1038/s41598-025-30991-6","workflowStages":[]},"version":"v1","identity":"rs-7424300","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7424300","identity":"rs-7424300","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00