Abstract
This study proposes additions and modifications to the Simple model for estimating Carbon Net Primary Productivity (CNPP) of agricultural crops, and parametrization and evaluation for the major grain and perennial forage crops in Brazil. CNPP was initially parameterized for soybean and maize using micrometeorological data, and for wheat, bean and perennial forage (Urochloa brizantha cv. Marandu) using agrometeorological experimental data and/or data from the literature. Subsequently, calibrations for reference cultivars were performed, grouping cultivars by crop phenological characteristics and edaphoclimatic regions using farm-level data. CNPP accurately simulated leaf area index, evapotranspiration, and biomass dry matter production and allocation for soybean and maize when evaluated at the sites with micrometeorological data (R² > 0.76, NSE > 0.56, and RRMSE < 38% for all variables). Simulations for wheat (R² = 0.2 and RRMSE = 31.8% for yield), common bean (R² = 0.01 and RRMSE = 48.5%) and perennial forage (R² = 0.33 and RRMSE = 27.2%) exhibited lower performance due to only yield data being available. Module performance was always lower for on-farm trials than for research site trials, due to lower quality environmental data and higher genotypic variability. Nonetheless, the resulting statistics (RRMSE 0) support this module’s efficacy in predicting crop productivity in major Brazilian agricultural areas, with statistics similar to those obtained by others. By employing a reduced and efficient parameter set, the CNPP module achieves enhanced performance and enables robust calibration across diverse crops, genotypes, and management schemes in multiple locations.
CORE IDEAS
•
CNPP is a simple and calibratable approach to estimate crop biomass.
•
The module can be coupled with PROCS to estimate carbon sequestration.
•
CNPP is a valuable tool for carbon sequestration in Brazilian agriculture.
TITLE: Calibration and evaluation of a Carbon Net Primary Productivity (CNPP) module based on the Simple model
Abstract
This study proposes additions and modifications to the Simple model for estimating Carbon Net Primary Productivity (CNPP) of agricultural crops, and parametrization and evaluation for the major grain and perennial forage crops in Brazil. CNPP was initially parameterized for soybean and maize using micrometeorological data, and for wheat, bean and perennial forage ( Urochloa brizantha cv. Marandu) using agrometeorological experimental data and/or data from the literature. Subsequently, calibrations for reference cultivars were performed, grouping cultivars by crop phenological characteristics and edaphoclimatic regions using farm-level data. CNPP accurately simulated leaf area index, evapotranspiration, and biomass dry matter production and allocation for soybean and maize when evaluated at the sites with micrometeorological data (R² > 0.76, NSE > 0.56, and RRMSE < 38% for all variables). Simulations for wheat (R² = 0.2 and RRMSE = 31.8% for yield), common bean (R² = 0.01 and RRMSE = 48.5%) and perennial forage (R² = 0.33 and RRMSE = 27.2%) exhibited lower performance due to only yield data being available. Module performance was always lower for on-farm trials than for research site trials, due to lower quality environmental data and higher genotypic variability. Nonetheless, the resulting statistics (RRMSE 0) support this module’s efficacy in predicting crop productivity in major Brazilian agricultural areas, with statistics similar to those obtained by others. By employing a reduced and efficient parameter set, the CNPP module achieves enhanced performance and enables robust calibration across diverse crops, genotypes, and management schemes in multiple locations.
Key words : carbon estimation, biomass production, crop model, carbon net primary productivity.
Introduction
Increasing computational processing capacity and availability of experimental datasets worldwide (https://agmip.org/) have facilitated the development and use of increasingly complex models (e.g., DSSAT and APSIM). The simultaneous emergence of simpler, less glamorous models (e.g., SIMPLE and SSM-iCrop2) with reduced process and parameter sets is an important parallel development, as their easier parametrization and implementation suits them for cases where data are scarce or incomplete.
Most biophysical crop models are designed to estimate biomass and yield, but even within that common mission their scope varies considerably. For example, DSSAT aimed to integrate knowledge about soil, climate, crops, and management to support better decisions about transferring production technology from one location to another (Jones et al., 2003). In contrast, APSIM was designed for accurate predictions of crop production in relation to climate, genotype, soil and management factors, in order to optimize resources and address management issues in farming systems (Keating et al., 2003). Meanwhile, simpler models such as SIMPLE (Zhao et al., 2019) are more flexible in how they are implemented and calibrated, as they are used for less-studied, less commodified crops.
Regardless of the purpose for which a crop model was developed, the growing demand for climate change policies has recently concentrated crop modeling efforts on understanding crops’ carbon emission and/or sequestration. Integrating this scientific knowledge to meet the needs of the carbon market is ongoing and incomplete. Intellectual property issues such as patents, proprietary or unavailable source code, and conflicts of interest currently limit the use or adaptation of many existing models for commercial purposes. It is therefore both important and challenging to develop new crop models that can estimate dry matter production and carbon inputs to soil carbon models, across many crops, environments, and management scenarios, for use by farmers, consultants, and policy makers.
The CNPP module is based on the SIMPLE model equations (Zhao et al., 2019), extended to simulate total aboveground and belowground biomass production and their components, for major tropical crops. The CNPP simulates aboveground and belowground biomass production separately, and also distinguishes between their harvested and residual components. This “accounting” enables coupling the CNPP module to a soil carbon stock model, providing the residual biomass both during the crop cycle and after harvest events. The development philosophy emphasizes an implementation with a reduced set of equations and parameters, facilitating calibration across diverse crops, genotypes, environments, and management practices.
The CNPP was implemented as a module integrated within a broader simulation platform, the Conecta Pro Carbono (https://procarbonoconecta.bayer.com/), specifically coupled to the ProCarbon-Soil (PROCS) soil carbon stock model (Barioni et al., 2025). This framework was designed for commercial applications, specifically for estimating soil carbon stock within agricultural environments across various crop types, rotations, and management practices. The present study aims to: (i) describe the CNPP module processes and parameters; (ii) parameterize and evaluate the module for different crop types including perennial forage; and (iii) evaluate the module for simulating these crops in Brazilian plantations.
Methods
Module description
The CNPP is a crop module that simulates crop growth, biomass production, and harvested and residual biomass in a daily time step. It was based on the SIMPLE model (Zhao et al., 2019), particularly in its foundational equations for phenology, growth, and the influence of solar, temperature and water on plant development. The CNPP’s phenological framework, including harvest timing, aligns with SIMPLE by triggering events based on accumulated thermal time (TT reaching a crop-specific target), an approach also employed by CERES-Wheat (Ritchie et al., 1985). However, CNPP diverges in its soil water balance calculation and water stress functions (Section 2.11). Furthermore, modifications to biomass production are also incorporated (Section 2.12). The subsequent sections will describe the equations for CNPP’s major processes that differ from those in SIMPLE.
2.1.1 Soil Water Balance
The root zone depth ( RZD, mm day -1 ) for a given date is calculated using Eq. 1. The active zone depth ( AZD, mm) is closely related to RZD, as described in Eq. 2.
Eq. 1 \(\text{RZD}_{t}=\left\{\par\begin{matrix}RZDGR,\ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ t=0\\ \min\left(\left(\text{RZD}_{t-1}+\min\left(\left(0.1+{f\left(\text{Solar}\right)}_{t-1}\right),\ 1\right)\ \times\ RZDGR\right),\ RZDM\right),\ \ t>0\\ \end{matrix}\right.\ \)
Eq. 2 \(\text{AZD}_{t}=\ \left\{\par\begin{matrix}\text{AZD}_{i},\ \ \ \ \ \ t=1\\ \text{RZD}_{t},\ {\text{\ \ \ \ \ }\text{AZD}}_{t-1}\leq\ \text{RZD}_{t}\\ \text{AZD}_{t-1},\ \ \ \ \ \ \text{AZD}_{t-1}>\ \text{RZD}_{t}\\ \end{matrix}\right.\ \)
where RZDGR is the root zone daily growth rate (mm); RZD t – 1 is the root zone depth on day t–1, RZDM is the maximum root zone depth (mm), f(Solar) t – 1 is fraction of solar radiation intercepted on day t–1 (dimensionless), and AZD t is the active zone depth (mm) on day
t, with initial value AZD i .
Soil water is computed daily for two different soil layers: (i) the near-surface soil water active zone ( SWAZ ), and (ii) the deeper soil water passive zone ( SWPZ ). The upper layer can “gain” water simply by the root depth increasing; this water transfer ( WT ) is handled as follows:
Eq. 3\(\text{WT}_{t}=\ \frac{\text{AZD}_{t}\ -\ \text{AZD}_{t-1}}{DMX\ -\ \text{AZD}_{t-1}}\ \times\ \text{SWPZ}_{t-1}\ \)
Eq. 4 \(\text{SWAZ}_{t}=\left\{\par\begin{matrix}\text{AZD}_{t}\ \times\ AWC\ \times\ 0.7\ +\ ERAIN-\ \text{ET}_{t},\ \ t=0\\ ERAIN\ +\text{SWAZ}_{t-1}\ +\ \text{WT}_{t}-\ \text{ET}_{t},\ \ t>0\\ \end{matrix}\right.\ \)
Eq. 5 \(\text{SWPZ}_{t}=\left\{\par\begin{matrix}{(DMX-\ AZD}_{t})\ \times\ AWC\ \times\ 0.7-\ \text{DR}_{t}\,\ \ t=0\\ ERAIN\ +\text{SWAZ}_{t-1}-\ \text{WT}_{t}-\ \text{DR}_{t},\ \ t>0\\ \end{matrix}\right.\ \)
where DMX is the maximum soil depth (mm); AWC is the fraction of soil water that is plant-available (dimensionless); ERAIN is effective rainfall (precipitation that infiltrates into the soil on day t, mm); ET t is the evapotranspiration (mm) on day t ; DR t is deep drainage (mm) on day t, and DMX is 300 mm deeper than RZDM, providing a soil buffer zone below the maximum root zone depth.
When SWAZ exceeds the water-holding capacity of the AZD (i.e., SWAZ t > AWC × AZD ) the excess water is transferred to SWPZ . Likewise, when SWPZ exceeds its water-holding capacity, the excess water is lost as deep drainage. Water dynamics in the soil layers are proportional to the relative soil water gradient ( ∆ W , Eq.6) (mm), which regulates the water transfer between the SWPZ and SWAZ (Eq. 7, 8 and 9):
Eq. 6\(\mathrm{\Delta}_{W}\ =\frac{\text{SWPZ}_{t}}{\left(DMX-\text{AZD}_{t}\right)\ \times\text{\ AWC}}\ -\frac{\text{SWAZ}_{t}}{\text{AZD}_{t}\ \times\text{\ AWC}}\)
Eq. 7 \(\text{UF}_{t}=\ \left\{\par\begin{matrix}\mathrm{\Delta}_{W}\ \times\ \text{SWPZ}_{t}\ \times\ \text{TW}\ \times\frac{\min\left(\left(DMX\ -\ \text{AZD}_{t}\right),\ \text{\ AZD}_{t}\right)}{\max\left(\left(DMX\ -\text{\ AZD}_{t}\right),\text{\ AZD}_{t}\right)},\ \ \ \mathrm{\Delta}_{W}\geq 0\\ \mathrm{\Delta}_{W}\ \times\ \text{SWAZ}_{t}\ \times\ \text{TW}\ \times\frac{\min\left(\left(DMX\ -\text{\ AZD}_{t}\right),\ \ \text{\ AZD}_{t}\right)}{\max\left(\left(DMX\ -\ \text{AZD}_{t}\right),\text{\ AZD}_{t}\right)},\ \ \ \mathrm{\Delta}_{W}<0\\ \end{matrix}\right.\ \)
Eq. 8 \(\text{SWAZ}_{t}=\ \text{SWAZ}_{t}+\ \text{UF}_{t}\)
Eq. 9 \(\text{SWPZ}_{t}=\ \text{SWPZ}_{t}-\ \text{UF}_{t}\)
where TW is a factor of time (in days) to reach an equilibrium between the two layers, and UF t is the unsaturated flux between the two layers (mm).
When the water supply to the plant is low, then evapotranspiration ( ET ) is reduced. We define a crop-specific fraction of plant-available water, dryStartET, below which the crop is drought-stressed. From this fraction, we define a soil water stress factor DrystressET (Eq. 10) which is then used to calculate the crop’s reduced ET (Eq. 11).
Eq. 10\(\text{Drystres}s_{\text{ET}}=\ \max\left(\min\left(\frac{\text{SWAZ}_{t}\ \times\ \text{dryStartET}}{\text{AWC\ }\times\ \text{AZD}_{t}},1\right),0\right)\)
Eq. 11\(\text{ET}_{t}=\min\left(\text{Drystres}s_{\text{ET}}\times\ K_{C}\times\ \text{ETo}_{t},\ \text{SWAZ}_{t}\right)\)
where \(\text{ETo}_{t}\) is the reference evapotranspiration on day t (mm day -1 ), and Kc is the evapotranspiration coefficient of the current crop (dimensionless), which changes with growth stage.
Finally, excess water on the surface is calculated using the curve number method (USA, 1986), in which a coefficient is generated to calculate the surface runoff water, i.e., water that arrives at the soil surface through precipitation or irrigation but does not infiltrate.
2.1.2 Crop Growth
The CNPP’s plant growth is regulated by light, temperature, and water stress, as observed in equation Eq. 12. Daily aboveground biomass production (Eq. 12) has modifications in comparison to the equations proposed by Zhao et al. (2019) in the SIMPLE model, where the AGB has an additional impact of atmospheric CO 2 concentration ( f(Co 2 )) and heat ( f(Heat) ) functions. To the AGB and yield components calculated by SIMPLE, we add computation of belowground biomass as follows:
Eq. 12\(\text{AGB}_{t}=\text{ABG}_{t-1}+\left(\text{RUE}\ \times\ f(\text{Solar})\times\ \text{SRAD}\ \times\ f\left(\text{Temp}\right)\ \times\ f\left(\text{Water}\ \right)\times 10\right)\)
Eq. 13 \(\text{BHARV}_{t}=\ \text{AGB}_{t}\ \times\ HI\)
Eq. 14\(\text{BLIT}_{t}=\ \text{AGB}_{t}\ \times\ \left(1-\ HI\right)\)
Eq. 15 \(\text{BGB}_{t}=\ \text{AGB}_{t}\ \times\ RTS\)
where RUE is the radiation use efficiency (g DM MJ -1 ), f(Solar) is the fraction of solar radiation interception (dimensionless), SRAD is the daily solar radiation (MJ m -2 day -1 ), f(Temp) is a function (Eq. 16) describing how RUE responds to daily air temperature (g MJ -1 ), f(Water) is the function (Eq. 17) of RUE response to soil available water status, HI is the harvest index (dimensionless ratio of harvested mass to pre-harvest aboveground biomass), AGB t is the total aboveground biomass (kg DM ha -1 ), BHARV t is the aboveground harvested biomass (kg DM ha -1 ), BLIT t is the aboveground non-harvested biomass (added to litter after harvest events, kg DM ha -1 ), BGB t is the belowground biomass (kg DM ha -1 ), and RTS is the root to shoot factor that calculate the root’s biomass in relation to aboveground biomass (dimensionless).
The f(Temp) function (Eq. 16) is a trapezoidal function that adjusts RUE given the daily mean air temperature T mean and four reference temperatures (°C): T base is the basal temperature below which daily biomass production is zero; T opt1 and T opt2 are respectively the lower and upper ends of the optimal temperature range within which daily biomass production is maximized; and T ext is the extreme (high) temperature above which daily biomass production is zero.
Eq. 16 \(f\left(\text{Temp}\right)=\left\{\par\begin{matrix}\max\left(\frac{T_{\text{mean}}-\ T_{\text{base}}}{T_{\text{opt}1}-\ T_{\text{base}}},\ 0\right),\ \ \ T_{\text{mean}}<\ T_{\text{opt}1}\ \\ 1,\ \ \ T_{\text{opt}1}<T_{\text{mean}}\ T_{\text{opt}2}\\ \end{matrix}\right.\ \)
The f(Water) is a dry stress factor determining the soil water status in relation to that fraction of plant available water at which stress starts to reduce plant growth (eq. 17).
Eq. 17\(f\left(\text{Water}\right)=\ \min\left(\frac{SWAZ\ \times\ \text{dryStartGR}}{\left({\text{DMX}-\ \text{AZD}}_{t}\right)\ \times\text{AWC}},\ 1\right)\)
where dryStartGR is the fraction of plant available water below which drought stress reduces plant growth (dimensionless).
The f(Solar) is regulated by the maximum intercepted solar radiation in the period of the maximum vegetative growth (eq. 18). Differently from SIMPLE, where the parameters I 50A and I 50 B are daily impacted by f(Heat) (a factor associated to heat stress) , in the CNPP the parameters are impacted by f(Water), with I 50 A being applied during the first half of the plant’s lifespan (TT t \(\leq\) T sum /2), and I 50B during the second:
Eq. 18\(f\left(\text{Solar}\right)=\ \min\left(\frac{\text{maxIntercept}}{1+\ e^{\left(-0.01\times\ \left(\text{TT}_{t}-I50A\right)\right)}},\ \ \frac{\text{maxIntercept}}{1+\ e^{\left(0.01\ \times\ \left(\text{TT}_{t}-\left(T_{\text{sum}}-\ I50B\right)\right)\right)}},\ \ 1\right)\)
Eq. 19\(I_{50A,t}=\ I_{50A,t-1}+I50maxW\ \times\ \left(1-\ f\left(\text{Water}\right)\right),\ \ \ \text{TT}_{t}\leq\ \frac{T_{\text{sum}}}{2}\)
Eq. 20\(I_{50B,t}=I_{50B,t-1}+I50maxW\ \times\ \left(1-\ f\left(\text{Water}\right)\right),\ \ \ \text{TT}_{t}>\ \frac{T_{\text{sum}}}{2}\)
where maxIntercept is the maximum fraction of solar radiation interception (dimensionless), I 50A,t is the thermal time requirement (ºC period -1 ) after sowing fraction of light interception to reach 50%, at day t, I 50B,t is the thermal time requirement (ºC period –1 ) from maturity backwards for light interception to reach 50% on day t, I50maxW is the maximum daily increase in I50A and I50B due to water stress (ºC) , TT t is the thermal time (ºC) at day t, and T sum is the total thermal time required from sowing to physiological maturity (ºC lifespan –1 ).
Field dataset descriptions
CNPP calibration and evaluation were performed for four annual crops and one perennial forage species: soybean [ Glycine max (L.) Merr.], maize ( Zea mays L. ), wheat ( Triticum aestivum L.), common bean ( Phaseolus vulgaris L.), and perennial forage [ Urochloa (syn. Brachiaria ) brizantha (Hochst ex A. Rich.) Stapf cv. Marandu]. Soybean and maize data were based on data collected in a well-monitored micrometeorological experiment (dataset 1). Wheat, common bean, and “Marandu” palisade grass data were collected in agrometeorological experiments, datasets 2, 3, and 4, respectively (Table 1). Additional datasets from a breeding program, representing on-farm conditions (datasets 5, 6, 7, and 8 respectively for soybean, maize, wheat, and common bean), were distributed across the most important Brazilian agricultural areas (Fig. 1). For these datasets, only grain yield data were available; no information was given for other biomass components such as aboveground residue.
Table 1. Summary of the datasets used for the parameterization, calibration, and evaluation of the Crop Net Primary Productivity (CNPP).
| 1 | Soybean | 41°9’53”N, 96°28’12”W (US-Ne2) & 41°10’46” N, 96°26’22” W (US-Ne3) | Dfa | LAI, ET, Biomass |
| Maize | ||||
| 2 | Wheat | 27°54’22” – 30°20’ 27”S, 51°1’17” – 54°53’17”W | Cfa, Cfb | Yield |
| 3 | Common bean | 9°28’2” – 28°30’33”S, 35°49’43” – 55°48’0”W | As, Aw, Cwa, Cfa, Cwb, Cfb | Yield |
| 4 | Perennial forage | 22°42’0”S, 47°30’0”W | Cwa | Biomass |
| 5 | Soybean | 3°3’49” – 33°11’47”S, 44°7’21” – 62°57’41”W | Af, Am, Aw, Cfa, Cfb, Cwa, Cwb | Yield |
| 6 | Maize | 1°46’39” – 33°11’52”S, 36°48’7” – 60°51’40”W | Af, Am, As, Aw, Cwa, Cwb, Cfa, Cfb | Yield |
| 7 | Wheat | 12°8’38” – 30°20’21”S, 43°46’13” – 55°47’42”W | Cwa, Cfa, Cwb, Cfb, Aw | Yield |
| 8 | Common bean | 7°30’0” – 29°11’60”S, 43°45’ 35” – 57°24’47”W | Am, Aw, Cwa, Cfa, Cwb, Cfa | Yield |
Fig. 1. Agrometeorological sites and on-farm plots used for parametrizing and evaluating the CNPP in Brazil. Köppen climate classification map is based on Alvares et al. (2013).
Dataset 1 is part of a long-term field experiment conducted at the Agricultural Research and Development Center (ARDC) of the University of Nebraska-Lincoln (UNL) near Mead, NE, USA, which has long-term hourly measurements of surface fluxes using via eddy covariance, soil moisture monitoring, crop phenology, and biomass time-series observations. Two sites were used: (i) US-Ne2 (41°9’53”N, 96°28’12”W, elevation 362 m, mean average temperature 10.08 ºC, and mean annual precipitation 788.89 mm), and (ii) US-Ne3 (41°10’46” N, 96°26’22” W, elevation 363 m, mean average temperature 10.11 ºC, and mean annual precipitation 783.68 mm). The US-Ne2 site is irrigated with a center pivot system, while US-Ne3 relies on rainfall (Suyker, 2024a, 2024b). The climate is designated Dfa (Humid Continental: humid with severe winter, no dry season, hot summer) in the Köppen climate classification system. The soils are deep silty clay loams (Mollic Hapludalfs, Pachic Argialbolls, and Vertic Argialbolls). Data from both sites (US-Ne2 and US-Ne3) are based on a maize-soybean crop rotation system under no-till since 2001 (details at https://ameriflux.lbl.gov/). For the present study, data from 2001 to 2014 were used, varying the crop and the irrigation scheme.
Dataset 2 is from wheat breeding trials, and includes information on grain yield, cultivar cycle, sowing and harvesting dates, and experimental location. This dataset includes non-irrigated experiments conducted during the off-season (winter season) between 2009 and 2022 at eight sites (Fig. 1): five in Brazil’s wheat producing region I (WPR I) (Cunha et al., 2006), and three in WPR II. A total of 132 cultivars, ranging from early to medium growth cycles, were evaluated in these experiments. Thus, dataset 2 comprises a total of 4,745 experiments: the calibration subset (totaling 2,948 experiments) includes data from three sites in WPR I and two in WPR II, while the evaluation subset (1,797 experiments) contains data from two sites in WPR I and one in WPR II.
Dataset 3 contains data on common bean obtained from the literature (Supplementary Material) from irrigated and non-irrigated experiments conducted at various times between 2000 and 2018. This dataset contains information on grain yield, cultivar cycle, sowing and harvesting dates, and experimental location. Dataset 3 includes a total of 71 cultivars from 23 sites (Brazilian municipalities; Fig. 1), covering growth habits I, II, and III. This dataset comprised a total of 161 experiments and was used only for module calibration.
Dataset 4 comprised data on the perennial forage “Marandu” palisade grass collected from an experimental area at the University of São Paulo, Luiz de Queiroz College of Agriculture in Piracicaba, São Paulo, Brazil (22°84’20”S, 47°83’00”W). The experiment was conducted between April 2011 and April 2013 under four treatments, the product of (a) irrigated versus non-irrigated, and (b) mechanized biomass harvesting every ~28 days versus every ~42 days. Further details on the experiment, data collection, and irrigation management can be found in Pequeno et al. (2014) and Pequeno et al. (2018). The main collected variables included herbage dry mass production (biomass harvested above a 10-cm stubble height), live perennial forage dry mass remaining (stubble dry mass), leaf area index (LAI), canopy light interception (LI), and harvest date. Weather data were obtained from a weather station located 1.8 km from the experimental area.
For soybean (dataset 5; breeding trial data), a split-block design was used, having 50,900 crop records (unique combination of genotype, cultivation year and site location) across 509 sites (municipalities), nested within 32 micro-regions. Data were collected from 2017 to 2024 across two growing seasons: rainy season, the most productive period; and off-season, a dry and less productive period. The dataset has 572 coded genotypes binned into five Relative Maturity (RM) groups, designated RM5 through RM9 . All fields have information on yield, sowing and harvest dates, with a subset of 14,373 records also having information on date of physiological maturity. This subset was used in the calibration procedure to precisely estimate soybean lifespan and harvest date. Further information is available in section 2.3, Parametrization and evaluation procedures .
Dataset 6 is structured similarly to dataset 5: its 24,139 maize records (unique combinations of genotype, cultivation year and site location), located in 433 sites, are binned into five mega environments (or breeding regions): “Subtropical Late”, “Subtropical Summer”, “Subtropical Winter”, “Tropical Highlands” and “Tropical Lowlands”. Data were collected from 2017 to 2024 across rainy season and off-season. A subset of 4,443 records also has information on silking date (the developmental stage when silks are first visible), which was used to estimate physiological maturity and harvest date. Further information is available in section 2.3, Parametrization and evaluation procedures .
Datasets 7 (260 wheat sites) and 8 (153 common bean sites), from a breeding program, were derived from rainfed farm fields conducted during the main and off-season between 2020 and 2023. The wheat sites were distributed across WPR I, II, III, and IV, and include cultivars ranging from early to medium growth cycles. Common bean sites were located in the main Brazilian bean producing regions (Fig. 1); however, information on cultivars and growth habits was not available. Datasets 7 and 8 respectively comprise 3,409 and 1,071 crop records.
Parametrization and evaluation procedures
For soybean and maize, dataset 1 has the most detailed measurements (“full parametrization”) including a flux tower, so it was used to parametrize the biophysical and physiological components of the module. Module parametrization was performed using a hierarchical structure of simulations wherein a combination of parameters varying in pre-defined intervals was tested. We followed the hierarchical sequence (i) radiation interception and absorption, (ii) physiological processes related to evapotranspiration and net primary production, and (iii) plant carbon allocation. Parametrization was conducted using the irrigated experiments, then evaluated using the rainfed experiments.
To parametrize soybean using dataset 5, specific parametrizations were performed using combinations of RM and Microregions, and the full parametrization was used as a reference. Parameters regulating (i) plant lifespan (cumulative temperature requirement from sowing to maturity); (ii) cumulative temperature requirement that regulates the plant light interception to reach 50%; (iii) cumulative temperature requirement that regulates the plant light interception to reach 50% due to leaf senescence; (iv) maximum light interception; (v) root to shoot; and (vi) harvest index (grain as a fraction of total aboveground biomass) were all recalibrated (Table 2). This dataset was partitioned into mutually exclusive calibration and evaluation subsets. The calibration set was used to develop and refine the approach, while the evaluation set was used for independent evaluation of its performance, avoiding data cross-contamination that would compromise validation.
Table 2. Summary of the crops and datasets used for the calibration and evaluation of the Crop Net Primary Production (CNPP). “C” – Calibration, “E” – Evaluation.
| Wheat | Bean | Perennial forage | Soybean | Maize | |
| Dataset 2 & 7 | Dataset 3 & 8 | Dataset 4 | Dataset 5 | Dataset 6 | |
| RZDM (mm) | C/E | C/E | - | - | - |
| Tbase (ºC) | C/E | - | - | - | - |
| Topt1 (ºC) | C/E | C/E | C | - | - |
| Topt2 (ºC) | C/E | C/E | C | - | - |
| ExtremeT (ºC) | C/E | C/E | C | - | - |
| RUE (g MJ -1 ) | C/E | C/E | C | - | - |
| Tsum (ºC cicle -1 ) | C/E | C/E | C | C/E | C/E |
| TbaseD (ºC cicle -1 ) | C/E | C/E | C | - | - |
| I50A (ºC period -1 ) | C/E | C/E | - | C/E | C/E |
| I50B (ºC period -1 ) | C/E | C/E | - | C/E | C/E |
| I50maxW (ºC day) | C/E | C/E | - | - | - |
| MaxIntercept | C/E | C/E | C | C/E | C/E |
| DryStartET | C/E | C/E | C | - | - |
| DryStartGR | C/E | C/E | C | - | - |
| Kcmin | C/E | C/E | C | - | - |
| Kcmax | C/E | C/E | C | - | - |
| RTSR | C/E | C/E | - | - | C/E |
| HI | C/E | C/E | - | C/E | C/E |
Similar calibration/evaluation procedures were performed for maize using dataset 6, where a specific parametrization was performed for each MG group in the mega environments. The full parametrization was used as reference for those parameters that regulate plant lifespan; parameters that were recalibrated regulated maximum light interception, water stress, and harvest index.
For wheat, common bean, and perennial forage, an initial parametrization was conducted based on the literature (see Supplementary Material). Subsequently, module calibration and evaluation were performed starting from the initial parametrization and working from the intermediate-quality datasets 2, 3, and 4. For perennial forage, a limited data was available so only calibration was performed (dataset 4).
Module calibration for wheat was carried out separately for each combination of cycle classification (early and medium growth cycles) and WPR (I and II), using the calibration set. This calibration involved adjusting the same six parameter groups used in the soybean recalibration with dataset 5. The calibrated module was then assessed against the evaluation set. After the initial calibration, a recalibration was performed for wheat using just dataset 7 (from a breeding program). This dataset was also divided into calibration and evaluation sets. The calibration set (1,654 crop records) included data from 129 sites, while the evaluation set (1,755 crop records) contained data from 131 sites, both distributed across all four WPRs. In this step, the parameters regulating plant lifespan and water stress, maximum light interception, and harvest index were recalibrated. This recalibration was performed separately for each combination of cycle classification and WPR, using the calibrated parameters of dataset 2 as a reference. In WPR III and IV, where dataset 2 data were unavailable, parameter recalibration was performed using the calibration set of dataset 7, then evaluated using the evaluation set. In contrast, for WPR I and II, only module validation was performed with the evaluation set of dataset 7, using the parameters calibrated based on dataset 2, as its experimental fields had a higher level of control than did dataset 7’s farm fields.
For the calibration of common bean, the full dataset 3 was used due to the scant data available for each growth habit. This calibration was conducted separately for each growth habit (I, II and III), and involved regulating the same parameters used in the soybean recalibration and the wheat calibration. Subsequently, a recalibration was performed using dataset 8 (from a breeding program). This dataset was divided into calibration (548 crop records and 76 sites) and evaluation (523 crop records and 77 sites) subsets. Similar to wheat, the parameters regulating plant lifespan, water stress, maximum light interception, and harvest index were recalibrated in this step.
Finally, for perennial forage, local module calibration was conducted using the full dataset 4. This calibration involved adjusting the parameters related to (i) plant lifespan (cumulative temperature requirement from sowing to maturity); (ii) cumulative temperature requirement that regulates the plant light interception to reach 50%; (iii) maximum light interception; (v) water stress; and (vi) optimal temperature for biomass production. To adapt the module to a perennial crop—starting a new cycle after each biomass harvest, rather than starting from sowing as grain crops—an artificial sowing date was set; this date was estimated to make the modeled f(Solar) be as close as possible to the f(Solar) estimated from the biomass observed immediately post-harvest.
To ensure robust module performance, the calibration and evaluation phases were conducted using independent datasets, with evaluation sites excluded from the calibration procedure. Furthermore, a temporal variability was incorporated to minimize biases associated with homogeneous climatic conditions.
Module Evaluation
We use the linear modeling capability of R (function “lm”; R Core Team, 2024) to evaluate the module’s performance in predicting crop yields. For each calibration or validation instance, we calculated the coefficient of determination (R 2 ; Eq. 20), root mean square error (RMSE; Eq. 21), relative root mean square error (RRMSE; Eq. 22), mean bias (MB; Eq. 23), and Nash-Sutcliffe Efficiency (NSE; Eq. 24):
Eq. 20\(R^{2}=1-\frac{\sum_{i=1}^{n}{(a_{i}-\ {\widehat{a}}_{i})^{2}}}{\sum_{i=1}^{n}{(a_{i}-\ {\overline{a}}_{i})^{2}}}\)
Eq. 21\(\text{RMSE}=\ \sqrt{\frac{\sum_{i=1}^{n}\left({\widehat{a}}_{i}-\ a_{i}\right)^{2}}{n}}\)
Eq. 22\(\text{RRMSE}=\frac{\sqrt{\frac{\sum_{i=1}^{n}\left({\widehat{a}}_{i}-\ a_{i}\right)^{2}}{n}}\ }{\frac{\sum_{i=1}^{n}{\widehat{a}}_{i}}{n}}\ x\ 100\)
Eq. 23\(\text{MB}=\ \sum_{i=1}^{n}\frac{\left({\widehat{a}}_{i}-\ a_{i}\right)}{n}\)
Eq. 24\(\text{NSE}=1-\frac{\sum_{i=1}^{n}{(a_{i}-\ {\widehat{a}}_{i})^{2}}}{\sum_{i=1}^{n}{(a_{i}-\ {\overline{a}}_{i})^{2}}}\)
where \(a_{i}\) denotes the observed values, \({\widehat{a}}_{i}\) the predicted values, \(\overline{a}\) the average of the observed values, and n the number of observations.
Results
Full CNPP parametrization (soybean and maize)
High performance, characterized by high R² values and low RMSE, RRMSE, and MB values (Table 3), was achieved in the simulations of LAI, evapotranspiration (ET), and aboveground biomass for soybean (Figs. 2, 3, and 4) and corn (Figs. 5, 6, and 7). Consistency between the calibration dataset (US-Ne2) and the evaluation dataset (US-Ne3) was evident for both crops, suggesting that the parameters derived from the calibration conducted on the irrigated site (US-Ne2 used for calibration) effectively predicted biomass on the rainfed site (US-Ne3 used for evaluation) (Figs. 4 and 7). We emphasize all analysis of LAI, ET and Biomass were performed base on the whole lifespan.
The CNPP demonstrated strong performance in estimating evapotranspiration (ET) for maize, with RRMSE values of 32.3% (irrigated site US-Ne2, calibration) and 38% (non-irrigated site US-Ne3, evaluation). These results position CNPP favorably among 42 models evaluated by Kimball et al. (2023), in which variations in modeled ET spanned 10–40% at 41–100 days after planting, and 10–70% for Phase 4 simulation, at the same sites. CNPP likewise exhibited superior ET estimation for soybean, yielding RRMSE values of 32.9% (irrigated US-Ne2) and 37.5% (non-irrigated US-Ne3). This contrasts with the Normalized Root Mean Square Error (nRMSE) range of 28.8% to over 120% observed across the soybean lifespan in 15 model variations or “flavor of models” (six APSIM and nine DSSAT-CSM-CROPGRO variants) in an in-season experiment (Silva et al., 2025).
Regarding LAI simulation in rainfed maize, Amiri et al. (2022) using CERES-Maize reported nRMSE values of 17–22% (calibration) and 9–17% (validation). CNPP significantly outperformed CERES-Maize, achieving RRMSE values of 4.3% (irrigated US-Ne2) and 4.6% (non-irrigated US-Ne3). For soybean LAI, Silva et al. (2025) reported an nRMSE range of 28.2% to over 40% across 15 model variants. In comparison, CNPP yielded RRMSE values of 5.4% (irrigated US-Ne2) and 7.7% (non-irrigated US-Ne3), indicating substantially better performance.
Finally, in simulating biomass of rainfed maize, Amiri et al. (2022) reported nRMSE values of 14–22% (calibration) and 5–16% (validation). CNPP again showed superior performance with RRMSE values of 3% (calibration US-Ne2) and 4.1% (validation US-Ne3). For soybean biomass, the nRMSE range across 15 models was 8.2% to over 40% (Silva et al., 2025). In contrast, CNPP achieved RRMSE values of 3.6% (irrigated US-Ne2) and 6% (non-irrigated US-Ne3), demonstrating a clear advantage over the other evaluated models.
The CNPP missed the late-season decline of soybean aboveground biomass (Fig. 4), which is due to leaf senescence and abscission following physiological maturity (R8 stage). The CNPP currently deposits leaves into the litter pool only at harvest. This module behavior is consistent with how maize’s (Fig. 7) senesced leaves typically remain attached to the main stem until harvest, resulting in minimal transfer of aboveground biomass to the litter pool even late in the growing season. At harvest, the CNPP allocates all non-harvested components (that is, all aboveground biomass except grain), including stems and leaves, to the litter pool, and roots directly to the soil (based on the species’ root-shoot parameters), providing input of organic carbon to the coupled soil biogeochemical model.
Table 3 Statistics (mean ± standard deviation) of CNPP simulations for aboveground biomass (kg ha -1 ), evapotranspiration (ET, mm day -1 ) and leaf area index (LAI, m 2 m -2 ) for maize and soybean in dataset 1. “ n ” is the number of harvest years.
| Soybean | Calibration | LAI (m 2 m -2 ) | 0.1 (±0.03) | 5.4 (±0.5) | 0.01 (±0.01) | 0.88 (±0.1) | 0.9 (±0.1) | 5 | 2 | |
| ET (mm day -1 ) | 0.8 (±0.1) | 32.9 (±3.8) | 0.2 (±0.1) | 0.75 (±0.1) | 0.8 (±0.03) | 5 | 3 | |||
| Biomass (kg ha -1 ) | 141.6 (±69.3) | 3.6 (±1.6) | 12.7 (±14.5) | 0.95 (±0) | 0.97 (±0.02) | 5 | 4 | |||
| Evaluation | LAI (m 2 m -2 ) | 0.2 (±0.1) | 7.7 (±2.9) | 0.02 (±0.01) | 0.56 (±0.4) | 0.9 (±0.05) | 7 | 2 | ||
| ET (mm day -1 ) | 0.9 (±0.2) | 37.5 (±10.8) | 0.2 (±0.1) | 0.67 (±0.2) | 0.76 (±0.1) | 7 | 3 | |||
| Biomass (kg ha -1 ) | 190.7 (±74.4) | 6 (±2.2) | 13 (±26.4) | 0.85 (±0.1) | 0.95 (±0.04) | 7 | 4 | |||
| Maize | Calibration | LAI (m 2 m -2 ) | 0.1 (±0.05) | 4.3 (±1.1) | 0.0002 (±0.02) | 0.89 (±0.1) | 0.94 (±0.05) | 9 | 5 | |
| ET (mm day -1 ) | 0.8 (±0.1) | 32.3 (±5.1) | 0.1 (±0.2) | 0.79 (±0.1) | 0.81 (±0.04) | 9 | 6 | |||
| Biomass (kg ha -1 ) | 297.3 (±112.6) | 3 (±1.3) | 30.6 (±40.2) | 0.97 (±0.02) | 0.99 (±0.01) | 9 | 7 | |||
| Evaluation | LAI (m 2 m -2 ) | 0.1 (±0.05) | 4.6 (±1.5) | 0.01 (±0.01) | 0.85 (±0.1) | 0.92 (±0.05) | 6 | 5 | ||
| ET (mm day -1 ) | 0.8 (±0.1) | 38 (±4.5) | 0.1 (±0.1) | 0.72 (±0.1) | 0.78 (±0.05) | 6 | 6 | |||
| Biomass (kg ha -1 ) | 333 (±84.5) | 4.1 (±1.2) | 23.1 (±35.9) | 0.93 (±0.02) | 0.98 (±0.01) | 6 | 7 |
RMSE: Root Mean Square Error, RRMSE: Relative Root Mean Square Error, MB: Mean Bias, NSE: Nash-Sutcliffe Efficiency, R 2 : coefficient of determination. Fig. 2. Simulated and observed soybean leaf area index (LAI) at the dataset 1 sites used for calibration (left) and evaluation (right). Fig. 3. Simulated and observed soybean evapotranspiration (ET) at the dataset 1 sites used for calibration (left) and evaluation (right). Red line is the 1:1 line. Fig. 4. Simulated and observed soybean aboveground biomass at the dataset 1 sites used for calibration (left) and evaluation (right). Fig. 5 Simulated and observed maize leaf area index (LAI) at the dataset 1 sites used for calibration (left) and evaluation (right). Fig. 6. Simulated and observed maize evapotranspiration (ET) at the dataset 1 sites used for calibration (left) and evaluation (right). Red line is the 1:1 line. Fig. 7. Simulated and observed maize aboveground biomass at the dataset 1 sites used for calibration (left) and evaluation (right).
CNPP parametrization for low-data crops (wheat, bean, and perennial forage)
While the parametrization for soybean and maize was based on micrometeorological measurements (dataset 1), the parametrization for wheat, common bean, and perennial forage were partially based on literature-derived parameter values. Moreover, the wheat (dataset 2) and common bean (dataset 3) module calibrations lacked long-term measurements and were limited to yield observations only. Field experiments had an additional uncertainty: breeding trials seldom collect local climate and soil data. Consequently, for these datasets the relationship between simulated and observed wheat and bean yields (Fig. 8) is noisier, as crucial variables such as LAI and ET could not be rigorously evaluated.
For wheat, a lower performance was noted (Table 4). These results are lower when comparable to those reported in published studies (R 2 values between 0.33 and 0.61) that utilized climate data at similar spatial resolutions (Richetti et al., 2024).
For perennial forage, the module exhibited low predictive accuracy for observed biomass yield (Fig. 9), under both irrigated and rainfed conditions. Values of 1338 of RMSE is very higher in comparison to the literature (Pequeno et al., 2014; Pequeno et al., (2018). However, improved predictions were observed when extended harvest intervals (Fig. 9C and D) were used in contrast to a shorter intervals (Fig. 9A and 9B), possibly because the early regrowth from 10-cm stubble was not accurately simulated by the module.
Table 4 Statistics of CNPP simulations of yield (kg ha -1 ) for wheat and common bean, and biomass yield for perennial forage in experimental sites (datasets 2, 3 and 4, respectively).
| Wheat | Calibration | 1453.4 | 30.5 | -382.1 | 0.14 | 0.2 |
| Evaluation | 1474.3 | 31.8 | -399.7 | 0.10 | 0.17 | |
| Bean | Calibration | 1029.4 | 48.5 | -318.7 | -0.81 | 0.01 |
| Perennial forage | Calibration | 1338.0 | 27.2 | -62.7 | -0.57 | 0.33 |
RRMSE: relative root mean square error, MB: mean bias, NSE: Nash-Sutcliffe efficiency, R 2 : coefficient of determination. Fig. 8. Simulated and observed grain yield of wheat and common bean. Red line is the 1:1 line.
Fig. 9 . Simulated (red line) and observed (circle) dry aboveground biomass production of irrigated (A and C) and non-irrigated (B and D) perennial forage harvested every ~28 (A and B) and ~42 days (C and D) during April 2011–April 2013 in Piracicaba, SP, Brazil, after calibration.
Evaluating local-scale yield simulations across Brazil’s agricultural areas
The modal cultivar calibrations—one calibration considering all cultivars per maturity group (MG) for soybean, maize, wheat, and common bean—were performed using a regionalized approach, one calibration combining all locations per region (Fig. 1). Module performance, however, was evaluated at the local scale: one average yield per MG. Significant variability was observed in the relationship between simulated and observed yield for both soybean and maize (Fig. 10), a consequence of the complex interaction of genotype, environment, and management practices that affects observed yield—factors not explicitly accounted for in the regional-scale calibration process. Furthermore, many soybean genotypes were grouped within the same GMR, and likewise many maize genotypes per MG. This aggregating inherently introduces uncertainty when applying regionally parameterized models to local scales, which are typically characterized by narrower ranges of genetics, environments, and management types.
Despite these challenges, the RRMSE and MB values (Table 5) were acceptable relative to other published studies. For example, Duarte and Sentelhas (2020) presented a similar analysis for 127 maize experimental sites, with calibration considering only flowering and harvest dates (calibration phase 1), similar to the calibration used for datasets 5, 6, 7 and 8. In that study, the RMSE values ranged from 2.12 to 5.73 Mg ha⁻¹, and R² ranged from 0.14 to 0.35 for the AEZ-FAO, APSIM-Maize, and DSSAT-CERES-Maize models. In contrast, the CNPP module achieved lower RMSE values of 1.99 Mg ha⁻¹ with an R² of 0.30 (calibration subset), and 1.96 Mg ha⁻¹ with an R² of 0.33 (evaluation subset), a clear advantage in predictive accuracy over the other models.
Battisti et al. (2017) also investigated the performance of five models—AEZ-FAO, AQUACROP, DSSAT, APSIM, and MONICA—across nine soybean experimental sites, with and without irrigation, over two years. Their calibration focused on life cycle phases (calibration phase 2), and had more detailed data than do datasets 5, 6, 7 and 8. Their reported RMSE values ranged from 0.44 to 1.958 Mg ha⁻¹, and R² values ranged from 0.39 to 0.79. Conversely, the CNPP module achieved RMSE values of 0.78 Mg ha⁻¹ with an R² of 0.14 (calibration subset) and 0.77 Mg ha⁻¹ with an R² of 0.08 (evaluation subset). Although the CNPP demonstrated to be comparable or superior predictive capability relative to the other models evaluated, a low values of R 2 show a moderate performance for some crops.
For common bean, the observed performance was significantly lower than that reported in the literature (Servín-Palestina et al., 2024). This disparity is attributed to limited site-specific information and the use of only yield data. Similarly, a low performance was noted for wheat when compared with the AquaCrop model’s results for different genotypes in southeastern Brazil (Rosa et al., 2019).
Table 5 Statistics of CNPP simulations for yield (Mg ha -1 ) for maize, soybean, wheat and common bean, based respectively on datasets 5, 6, 7, and 8.
| Soybean | Calibration | 0.78 | 24.84 | 0.04 | 0.12 | 0.14 |
| Evaluation | 0.77 | 23.53 | 0.08 | 0.03 | 0.08 | |
| Maize | Calibration | 1.99 | 28.78 | 0.06 | 0.29 | 0.3 |
| Evaluation | 1.96 | 28.59 | 0.14 | 0.31 | 0.32 | |
| Wheat | Calibration | 1.22 | 35.72 | 0.33 | -0.51 | 0.04 |
| Evaluation | 1.27 | 37.11 | 0.26 | -0.53 | 0.05 | |
| Bean | Calibration | 0.84 | 43.62 | 0.02 | -0.52 | 0.04 |
| Evaluation | 0.91 | 46.3 | 0.08 | -0.93 | 0.001 |
RRMSE: Relative Root Mean Square Error, MB: Mean Bias, NSE: Nash-Sutcliffe Efficiency, R 2 : coefficient of determination. Fig. 10. Simulated and observed yields (Mg ha -1 ) of maize, soybean, wheat and common bean. Red line is the 1:1 line.
Evaluating the regional-scale yield simulations for Brazil
In contrast to the previous analysis (Section 3.3), which evaluated module performance locally, the regionalized results (Table 6) represent mean yields derived from aggregating measurements within subsets of genotype (GMR for soybean, and MG for maize and wheat), regions (micro-regions for soybean and mega environments for maize), across crop seasons. Figure 11 illustrates these groupings, specifically showing mean yields for soybean (averaged across GMRs within micro-regions across crop seasons), maize, wheat (the latter two each averaged across MGs within regions across crop seasons), and common bean. As hypothesized, regional-scale simulations (Table 5 and Figure 11) demonstrated improved performance compared to location-specific analyses (Table 4 and Figure 10), since the regional averaging reduces the impact of localized variability. With the exceptions of the R² for soybean and wheat, and the NSE for wheat, the consistency observed between calibration and evaluation datasets suggests the effectiveness of the presented approach for estimating regional yield.
Table 6 Statistics of CNPP regionalized yield simulations (Mg ha -1 ) for maize, soybean, wheat and common bean from datasets 5, 6, 7 and 8, respectively.
| Soybean | Calibration | 0.59 | 18.37 | 0.07 | 0.23 | 0.24 |
| Evaluation | 0.64 | 19.76 | 0.09 | 0.02 | 0.09 | |
| Maize | Calibration | 0.78 | 11.31 | 0.03 | 0.72 | 0.72 |
| Evaluation | 0.94 | 13.7 | 0.04 | 0.64 | 0.65 | |
| Wheat | Calibration | 0.77 | 25.85 | 0.28 | -0.12 | 0.29 |
| Evaluation | 0.87 | 30.59 | -0.22 | -0.82 | 0.19 | |
| Bean | Calibration | 0.47 | 22.49 | 0.1 | -2.72 | 0.2 |
| Evaluation | 0.65 | 30.24 | 0.28 | -12.0 | 0.06 |
Fig. 11. Simulated and observed regionalized yield (Mg ha -1 ) for maize, soybean, wheat and common bean from datasets 5, 6, 7, and 8 respectively. Red line is the 1:1 line.
Discussion
The CNPP module was initially calibrated using a high-quality dataset, with flux tower measurements and in-season biomass data. The module shows consistent accuracy for simulating ET, LAI and aboveground biomass. Comparisons with published studies revealed that CNPP’s performance was either equivalent to or better than several models in the literature for soybean (Silva et al., 2025) and maize (Amiri et al., 2022; Kimball et al., 2023) in the same or similar experimental sites.
In contrast to the CNPP’s high performance for calibrating and validating to the full parametrization data (dataset 1), the calibration to Brazilian conditions resulted in decreased performance, attributable mainly to the inherent limitations of the available data, and to higher genotypic and environmental variability. Crop modeling studies typically utilize experimental sites characterized by extensive measurements (e.g., LAI, total aboveground biomass, specific leaf area) and well-known environmental conditions (e.g., climate and soil) (Battisti et al., 2017; Dias et al., 2023). Differently, the calibration process performed using on-farm data (datasets 5, 6, 7, and 8) largely had only yield data, with phenological information available only for a small subset. Furthermore, the soil and climate information was more at a regional than a local scale.
Despite these challenges, CNPP exhibited performance comparable to or exceeding those reported in the literature for studies employing similar calibration procedures (e.g., Battisti et al., 2017; Duarte and Sentelhas, 2020). Moreover, the analysis of on-farm data, characterized by an operational rather than a strictly experimental approach, suggested the module’s suitability for predicting crop yield at a regional scale. This regional predictive capacity is of considerable importance for integration with SOC models, indicating CNPP’s potential to provide reliable estimates of carbon input across the evaluated regions.
While parametrizing a crop model often relies on flux tower data, which provides high-resolution measurements of ecosystem exchanges, an alternative parametrization strategy must be employed when such data are unavailable. Unlike the case for maize and soybean, the species parameters for wheat, common bean, and perennial forage had to be calibrated without flux tower data. However, parameters governing species-specific processes were adjusted based on established literature values (see Supplementary Material). Despite the loss of accuracy compared to flux tower-assisted calibration, this approach offers significant practical advantages by eliminating the need for resource-intensive, long-term field studies, and has demonstrated effectiveness in previous studies (Soltani et al., 2020; Vergel et al., 2022; Zhao et al., 2019).
Wheat exhibits high phenotypic plasticity, growing in both subtropical and tropical regions under rainfed and irrigated conditions. In Brazil, some wheat cultivation occurs in tropical regions, including the Neotropical savanna located in the central part of the country, but production is concentrated in the southernmost regions, where the rainy season is more accentuated. While this region is generally favorable to crop production, the incidence of rain coinciding with physiological maturity can induce pre-harvest sprouting, resulting in yield reductions (Rigatti et al., 2019). This phenomenon, which is not currently represented in the module, complicates the calibration processes.
For beans, the data needed to calculate available water were frequently incomplete. Some sites provided only soil texture information, while others offered only site coordinates, necessitating the use of external databases to estimate soil properties and thence available water. Meanwhile, for perennial forage the primary challenge in module calibration concerned regrowth. In field conditions, regrowth is a regenerative process driven by leaf development, with root compartments remaining intact. Regrowth itself is an inherently difficult event to model, and becomes even more challenging when animal grazing events are incorporated (Bosi et al., 2020). Although this study did not explicitly simulate the effects of grazing, the module simulates each “virtual grazing event” as a harvest, analogous to other crops, where compartments are reset immediately post-event. Even so, the model did not perform as well as some studies in the literature (Pequeno et al., 2014; Pequeno et al., 2018). This simplification was an alternative to the absence of a dedicated grazing process, and even though it didn’t perform so well, it was considered acceptable due to the simplicity of the model.
Regardless, the module’s simplicity, simplified physiological processes, and limited parameter set all facilitate calibration. This allows for recalibration at various scales, including regional, site-specific, genotypic, or even management-specific levels. Therefore, CNPP may benefit in future from the adoption of new practices in crop modeling and calibration, such as (i) using data assimilation based on remote sensing (Jin et al., 2018), (ii) employing hybrid approaches with artificial intelligence (e.g., Chang et al., 2023; Shahhosseini et al., 2021), or (iii) using CNPP in an ensemble scheme (e.g., Battisti et al., 2017; Duarte and Sentelhas, 2020). These approaches generally result in higher model accuracy compared to the performance observed in on-farm data using datasets 5, 6, 7, and 8, which are mostly based on operational procedures of harvest.
Conclusions
The CNPP module provides a simple, functional, and readily calibratable approach to estimate crop biomass production. Its capacity to incorporate new crops suggests broad applicability for estimating carbon assimilation by crops, and its ability to be coupled with soil biogeochemical models recommends its use in estimating carbon sequestration in agricultural environments. Specifically, within the major Brazilian maize, soybean, wheat and common bean production regions considered in this study, CNPP is expected to be a valuable tool for estimating carbon sequestration when coupled with the PROCarbon Soil (PROCS; Barioni et al. 2025) soil biogeochemical model.
References
Alvares, C. A., Stape, J. L., Sentelhas, P. C., Gonçalves, J. L. M., & Sparovek, G. (2013). Köppen’s climate classification map for Brazil. Meteorologische Zeitschrift, 22 (6), 711–728. https://doi.org/10.1127/0941-2948/2013/0507Amiri, E., Irmak, S., & Araji, H. A. (2022). Assessment of CERES-Maize model in simulating maize growth, yield and soil water content under rainfed, limited and full irrigation. Agricultural Water Management, 259, 107271. https://doi.org/10.1016/J.AGWAT.2021.107271Balboa, G. R., Archontoulis, S. V., Salvagiotti, F., Garcia, F. O., Stewart, W. M., Francisco, E., Prasad, P. V. V., & Ciampitti, I. A. (2019). A systems-level yield gap assessment of maize-soybean rotation under high- and low-management inputs in the Western US Corn Belt using APSIM. Agricultural Systems, 174 (April 2018), 145–154. https://doi.org/10.1016/j.agsy.2019.04.008Barioni, L. G., Valladão, B. A., Mourão, V. H. M., Karatay, Y. N., Damian, J. M., Moreno, L. M., & Silva, R. O. (2025). Procarbon-Soil—Procs: aA dynamic soil carbon model for improved model-data compatibility in carbon farming. SSRN, Preprint. https://doi.org/10.2139/ssrn.4812266Battisti, R., Bender, F. D., & Sentelhas, P. C. (2019). Assessment of different gridded weather data for soybean yield simulations in Brazil. Theoretical and Applied Climatology, 135 (1–2), 237–247. https://doi.org/10.1007/s00704-018-2383-yBattisti, R., Sentelhas, P. C., & Boote, K. J. (2017). Inter-comparison of performance of soybean crop simulation models and their ensemble in southern Brazil. Field Crops Research, 200, 28–37. https://doi.org/10.1016/j.fcr.2016.10.004Bosi, C., Sentelhas, P. C., Huth, N. I., Pezzopane, J. R. M., Andreucci, M. P., & Santos, P. M. (2020). APSIM-Tropical Pasture: A model for simulating perennial tropical grass growth and its parameterisation for palisade grass (Brachiaria brizantha). Agricultural Systems, 184, 102917. https://doi.org/10.1016/j.agsy.2020.102917Chang, Y., Latham, J., Licht, M., & Wang, L. (2023). A data-driven crop model for maize yield prediction. Communications Biology, 6 (1), 439. https://doi.org/10.1038/s42003-023-04833-yCunha, G. R., Pires, J. L. F., Maluf, J. R. T., Pasinato, A., Caierão, E., Silva, M. S., Dotto, S. R., Campos, L. A. C., Felício, J. C., Castro, R. L., Marchioro, V., Riede, C. R., Rosa Filho, O., Tonon, V. D., & Svoboda, L. H. (2006). Regiões de adaptação para trigo no Brasil. Circular Técnica, 20, Embrapa Trigo, 1–35.Dias, H. B., Cuadra, S. V., Boote, K. J., Lamparelli, R. A. C., Figueiredo, G. K. D. A., Suyker, A. E., Magalhães, P. S. G., & Hoogenboom, G. (2023). Coupling the CSM-CROPGRO-Soybean crop model with the ECOSMOS Ecosystem Model – An evaluation with data from an AmeriFlux site. Agricultural and Forest Meteorology, 342 . https://doi.org/10.1016/j.agrformet.2023.109697Duarte, Y. C. N., & Sentelhas, P. C. (2020). Intercomparison and performance of maize crop models and their ensemble for yield simulations in Brazil. International Journal of Plant Production, 14 (1), 127–139. https://doi.org/10.1007/s42106-019-00073-5Jin, X., Kumar, L., Li, Z., Feng, H., Xu, X., Yang, G., & Wang, J. (2018). A review of data assimilation of remote sensing and crop models. European Journal of Agronomy, 92, 141–152. https://doi.org/10.1016/j.eja.2017.11.002Jones, J. W., Hoogenboom, G., Porter, C. H., Boote, K. J., Batchelor, W. D., Hunt, L. A., Wilkens, P. W., Singh, U., Gijsman, A. J., & Ritchie, J. T. (2003). The DSSAT cropping system model. European Journal of Agronomy, 18 (3–4), 235–265. https://doi.org/10.1016/S1161-0301(02)00107-7Keating, B. A., Carberry, P. S., Hammer, G. L., Probert, M. E., Robertson, M. J., Holzworth, D., Huth, N. I., Hargreaves, J. N. G., Meinke, H., Hochman, Z., McLean, G., Verburg, K., Snow, V., Dimes, J. P., Silburn, M., Wang, E., Brown, S., Bristow, K. L., Asseng, S., … Smith, C. J. (2003). An overview of APSIM, a model designed for farming systems simulation. European Journal of Agronomy, 18 (3–4), 267–288. https://doi.org/10.1016/S1161-0301(02)00108-9Kimball, B. A., Thorp, K. R., Boote, K. J., Stockle, C., Suyker, A. E., Evett, S. R., Brauer, D. K., Coyle, G. G., Copeland, K. S., Marek, G. W., Colaizzi, P. D., Acutis, M., Alimagham, S., Archontoulis, S., Babacar, F., Barcza, Z., Basso, B., Bertuzzi, P., Constantin, J., … Zhou, W. (2023). Simulation of evapotranspiration and yield of maize: An inter-comparison among 41 maize models. Agricultural and Forest Meteorology, 333, 109396. https://doi.org/10.1016/J.AGRFORMET.2023.109396Pequeno, D. N. L., Pedreira, C. G. S., & Boote, K. J. (2014). Simulating forage production of Marandu palisade grass (Brachiaria brizantha) with the CROPGRO-Perennial Forage model. Crop and Pasture Science, 65 (12), 1335–1348. https://doi.org/10.1071/CP14058Pequeno, D. N. L., Pedreira, C. G. S., Boote, K. J., Alderman, P. D., & Faria, A. F. G. (2018). Species-genotypic parameters of the CROPGRO Perennial Forage Model: Implications for comparison of three tropical pasture grasses. Grass and Forage Science, 73 (2), 440–455. https://doi.org/10.1111/gfs.12329R Core Team. (2024). R: A language and environment for statistical computing (4.3.3). R Foundation for Statistical Computing. https://www.r-project.org/Richetti, J., Lawes, R. A., Whan, A., Gaydon, D. S., & Thorburn, P. J. (2024). How well does APSIM NextGen simulate wheat yields across Australia using gridded input data? Validating Continental-Scale Crop Model Simulations. European Journal of Agronomy, 158, 127212. https://doi.org/10.1016/j.eja.2024.127212Rigatti, A., Meira, D., Olivoto, T., Meier, C., Nardino, M., Lunkes, A., Klein, L. A., Fassini, F., Moro, É. D., Marchioro, V. S., & Souza, V. Q. (2019). Grain yield and its associations with pre-harvest sprouting in wheat. Journal of Agricultural Science, 11 (4), 142. https://doi.org/10.5539/jas.v11n4p142Ritchie, J. T., Godwin, D. C., & Otter-Nacke, S. (1985). CERES-wheat: a user-oriented wheat yield model . Preliminary Documentation. AGRISTARS Publication No. YM-U3- 04442-JSC-18892. Michigan State University.Rosa, S. L. K., Souza, J. L. M., & Tsukahara, R. Y. (2019). Performance of the AquaCrop model for the wheat crop in the subtropical zone in Southern Brazil. Pesquisa Agropecuária Brasileira, 55, e01238. https://doi.org/10.1590/S1678-3921.PAB2020.V55.01238Servín-Palestina, M., López-Cruz, I., Zegbe, J. A., Ruiz-García, A., Salazar-Moreno, R., & Cid-Ríos, J. Á. (2024). Calibration and evaluation of the SIMPLE crop growth model applied to the common bean under irrigation. Agronomy, 14, 917. https://doi.org/10.3390/agronomy14050917Shahhosseini, M., Hu, G., Huber, I., & Archontoulis, S. V. (2021). Coupling machine learning and crop modeling improves crop yield prediction in the US Corn Belt. Scientific Reports, 11 (1), 1606. https://doi.org/10.1038/s41598-020-80820-1Silva, E. H. F. M., Kothari, K., Pattey, E., Battisti, R., Boote, K. J., Archontoulis, S. V., Cuadra, S. V., Faye, B., Grant, B., Hoogenboom, G., Jing, Q., Marin, F. R., Nendel, C., Qian, B., Smith, W., Srivastava, A. K., Thorp, K. R., Vieira Junior, N. A., & Salmerón, M. (2025). Inter-comparison of soybean models for the simulation of evapotranspiration in a humid continental climate. Agricultural and Forest Meteorology, 365, 110463. https://doi.org/10.1016/J.AGRFORMET.2025.110463Soltani, A., Alimagham, S. M., Nehbandani, A., Torabi, B., Zeinali, E., Dadrasi, A., Zand, E., Ghassemi, S., Pourshirazi, S., Alasti, O., Hosseini, R. S., Zahed, M., Arabameri, R., Mohammadzadeh, Z., Rahban, S., Kamari, H., Fayazi, H., Mohammadi, S., Keramat, S., … Sinclair, T. R. (2020). SSM-iCrop2: A simple model for diverse crop species over large areas. Agricultural Systems, 182, 102855. https://doi.org/10.1016/j.agsy.2020.102855Suyker, A. (2024a). AmeriFlux BASE US-Ne2 Mead - irrigated maize-soybean rotation site (18–5). AmeriFlux AMP, (Dataset). https://doi.org/10.17190/AMF/1246085Suyker, A. (2024b). AmeriFlux BASE US-Ne3 Mead - rainfed maize-soybean rotation site (18–5). AmeriFlux AMP, (Dataset). https://doi.org/10.17190/AMF/1246086USDA. (1986). Urban hydrology for small watersheds. In Technical Release 55 (2nd ed.). United States Department of Agriculture. https://www.nrc.gov/docs/ML1421/ML14219A437.pdfVergel, A. P. R., Rodriguez, A. V. C., Ramirez, O. D., Velilla, P. A. A., & Gallego, A. M. (2022). A crop modelling strategy to improve cacao quality and productivity. Plants, 11, 157. https://doi.org/10.3390/plants11020157Zhao, C., Liu, B., Xiao, L., Hoogenboom, G., Boote, K. J., Kassie, B. T., Pavan, W., Shelia, V., Kim, K. S., Hernandez-Ochoa, I. M., Wallach, D., Porter, C. H., Stockle, C. O., Zhu, Y., & Asseng, S. (2019). A SIMPLE crop model. European Journal of Agronomy, 104 (August 2018), 97–106. https://doi.org/10.1016/j.eja.2019.01.009
Information & Authors
Information
Version history
Peer review timeline
Published
Agronomy Journal
Version of Record26 Mar 2026Published
Copyright
This work is licensed under a Non Exclusive No Reuse License.
Keywords
Authors
Metrics & Citations
Metrics
Article Usage
213views
118downloads
Citations
Download citation
Michel Colmanetti, Henrique Oldoni, Leandro Silva, et al.
Calibration and evaluation of a Carbon Net Primary Productivity (CNPP) module based on the Simple model. Authorea. 24 November 2025.
DOI: https://doi.org/10.22541/au.176402479.99853267/v1
DOI: https://doi.org/10.22541/au.176402479.99853267/v1
If you have the appropriate software installed, you can download article citation data to the citation manager of your choice. Simply select your manager software from the list below and click Download.
For more information or tips please see 'Downloading to a citation manager' in the Help menu.
Cited by
- ProCarbon‐Soil: A dynamic model for improved model‐data compatibility in carbon farming, Soil Science Society of America Journal, 90, 3, (2026).https://doi.org/10.1002/saj2.70218
Loading...
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.