
The Gold Fyfe dataset is the largest collection of daily chromatograms of one plant (n = 29 979, covering over 27 years). Despite its high chronobiological relevance, it has so far been only partially explored. In this study, we tested a new characterization method of the chromatograms, based on qualitative features perceptible to human observers (i.e. gestalt recognition through kinaesthetic engagement) using deep learning. A trained human evaluator identified three gestalts and labelled chromatograms accordingly. Three deep learning models (one learning-from-scratch and two transfer learning approaches) were developed and evaluated across three data partitions for training validity, performance, and reliability. An explainability analysis was performed on one model, and the models' descriptive ability was compared with traditional image analysis descriptors previously used. A preliminary chronobiological application was tested on the dataset (PSDs and cross-correlation analysis). Evaluator consistency was 100%. The transfer learning models outperformed the learning-from-scratch model across all evaluation phases. The explainability analysis confirmed that the models relied on the same morphological features identified by the evaluator. Few correlations with traditional descriptors were found, suggesting an expansion of the descriptive palette. The preliminary chronobiological analysis revealed a correlation between one class and seasonal weather parameters. The proposed deep learning models represent a first attempt to achieve a qualitative characterization of metabolomic fingerprinting images guided by human perception. The models successfully extended the descriptive palette of the Gold Fyfe dataset with human-relevant descriptors, and their preliminary chronobiological application yielded promising results, warranting further investigation.
Process-based models (PBMs) provide a mechanistic foundation for simulating complex genotype, environment, and management interactions across multiple scales. Integrating PBMs with agent-based models (ABMs), machine learning (ML), and advanced model coupling strategies (sequential, loose, tight, and framework-based) enables more comprehensive representations of agricultural systems. ABMs capture the adaptive behaviours of individual farmers, while ML techniques efficiently approximate nonlinear physiological processes or act as surrogates for computationally expensive submodules and reveal data-driven patterns, collectively enhancing the simulation of G × E × M interactions under variable climate scenarios. Digital twins (DTs), defined as systems that dynamically synchronize virtual models with real-time sensor data, extend coupled PBM–ABM–ML systems by enabling bidirectional feedback between computational models and farm operations. However, the implementation of DTs remains constrained by challenges such as data integration across heterogeneous sources, computational scalability, and semantic/technical interoperability making them appropriate only where continuous decision support and real-time feedback outweigh complexity and cost. This review evaluates the function of PBMs as the mechanistic core of DT architectures, surveys model integration techniques, and outlines the IT infrastructure required for operationalization. We conclude that digital twins are best deployed when real-time feedback is essential, whereas PBM–ABM–ML couplings suffice for research and policy applications requiring long-term scenario analysis.
Plant models are essential for understanding the mechanisms underlying complex plant processes and for predicting growth under varying environmental conditions. They play a central role in plant science and have direct applications in plant breeding, crop management, and related fields. Functional-structural plant models form a widely used class of models that explicitly represent plant structure. Because functional-structural plant models are difficult to implement from scratch, several frameworks have been developed to make them more accessible. However, none of the existing frameworks adopts an acausal modelling approach. Acausal modelling allows users to define systems as differential equation systems without prescribing causal relationships in advance, thereby improving model composability, reusability, and user-friendliness. We introduce PlantModules.jl, an acausal, differential-equation-based functional plant modelling framework designed to integrate with existing structural modelling frameworks for the creation of functional-structural plant models. The framework emphasizes modular model construction, extensibility and customizability, and provides core functionality centred on plant-water relations, offering a broadly applicable foundation for describing plant growth. We illustrate the framework's capabilities through three case studies, including validation against pine tree growth data. As an open-source Julia library, PlantModules.jl provides a flexible platform for future plant modelling efforts.
Limited wheat transpiration under rising vapour pressure deficit (VPD), i.e. VPD-sensitivity, has been shown to improve yields under terminal drought by conserving water. However, this trait may lead to yield penalties in well-watered environments. Using the US spring wheat belt as a case study, simulations were conducted to assess the effect of genotypic variation in transpiration VPD sensitivity on yield performance across the entire region relative to a genotype not expressing this trait. We used two crop models, APSIM and SSM-Wheat, representing different levels of algorithmic complexity, and historical and future climate scenarios for the 2050s and 2080s based on simulations using 20 global climate models across 457 locations. Regardless of the VPD threshold (1.5 kPa, VPD1.5 or 2 kPa, VPD2.0) at which the sensitivity was initiated, this trait led to systematic yield gains (13%-16% for VPD2.0, 40%-45% for VPD1.5) under historic climate scenarios across the entire region, regardless of the crop model, with high probability (> 0.8). Large yield benefits (16%-20% for VPD2.0, 29%-65% for VPD1.5) are also expected to occur throughout nearly all future climate scenarios, except for the most remote and extreme one (2080s_RCP8.5). Both crop models, however, diverged in identifying regions with the highest levels of yield increase. Overall, this study points to the key importance of introgressing transpiration sensitivity to VPD in US wheat germplasm and to the need for joint analyses of more than one crop model in predicting spatially resolved yield performance due to this trait.
Long-duration space missions will require the use of bioregenerative life support systems (BLiSS) to ensure renewability of the crews food and oxygen. Sustainable management of these systems will require accurate predictions of biomass accumulation and gas exchanges from crop production. Four energy cascade (EC) models have been developed to predict biomass yield and transpiration rates of BLiSS crop production. Lettuce cultivars Waldmann's Green and BG231-1251 were grown at a daily average temperature of 21 degrees C, ambient CO2, and a daily light integral of 22.03 mol m-2 d-1 to evaluate the ECs predictive abilities. All ECs overestimated lettuce biomass and underestimate transpiration with root mean square errors ranging between 49.58 to 96.99 g m-2 and 0.80 to 2.30 L m-2 d-1, respectively. Global sensitivity analysis indicated that the ECs biomass predictions are insensitive to temperature changes. The ECs are additive with environmental parameters, such as light and CO2, accounting for half of any output's sensitivity. Interactions between parameters accounted for, at most, a third of the ECs sensitivity. The ECs provide a baseline for BLiSS crop management; however, this comprehensive, retrospective analysis reveals systematic biases in predictions of lettuce crops, stemming from erroneous canopy closure methodology and lack of temperature responses. Further work on lettuce, or other crops, can use this analysis as a benchmark to improve the range, accuracy, and reliability of the ECs for sustaining BLiSS and the astronauts they support.
Noncoding regions mediate transcriptional adaptation to stress in plants, yet the genomic determinants driving gene expression responses remain poorly understood. Here, we evaluate the ability of two deep learning approaches, a convolutional neural network (CNN) and a transformer-based genomic language model pre-trained on plant genomes to predict changes in gene expression under various abiotic and biotic stress conditions. Using RNA-seq time-series data from Arabidopsis thaliana, we explored different strategies for summarizing expression dynamics to capture treatment-specific transcriptional changes relative to control conditions. Both models achieved low to moderate predictive performance, with pattern-triggered-immunity related treatments showing the strongest sequence-based predictability. While extending promoter regions upstream had a limited impact, including coding sequences significantly improved performance. Model interpretation revealed that the CNN recovered sequence features comparable to those identified by simple 6-mer based linear models, suggesting limited gains in regulatory insight from increased model complexity. These findings underscore both the promise and limitations of sequence-based models in uncovering the regulatory logic of induced plant stress responses.
Over the past decades, remote sensing has evolved beyond traditional site-based methods, enabling drought monitoring and enhancing our ability to manage and mitigate its impacts through agronomic management. However, at smaller scales, such as breeding or research plot trials, the timely assessment of crop water use remains a challenge due to spatial and temporal resolution limitations. In this study, we investigated the potential to integrate climatic/satellite data with field-deployed sensing systems in the estimation of wheat crop evapotranspiration across Australia. Using National Variety Trial (NVT) sites, infrared thermometers and RGB cameras were deployed in wheat plots to characterize canopy temperature fluctuations. In 41 wheat trials, we estimated evapotranspiration on a reference cultivar using a surface energy budget approach based on sensor data, resulting in similar estimates to those obtained from APSIM simulations (R2 = 0.89). We then used an expolinear modelling approach for the estimation of transpiration efficiency (TE) and evaporation from soil (ES) from the relationship between biomass and evapotranspiration when a limited number of observations are available. We showed the environmental variation in TE across sites was explained by nitrogen and water deficits, contributing to 20% of yield variability in NVT. We emphasized that this approach has strong potential for real-time monitoring of crop water use and TE traits, especially in trials with limited monitoring and sampling frequency. Overall, these results demonstrated the feasibility of estimating water use traits in the field using sensor-driven approaches and the utility of multi-site monitoring for characterizing temporal drivers of crop performance.
Decision support systems (DSS) in crop protection provide valuable support for pest risk prognosis and recommendations for pest control, enabling farmers to make better-informed decisions. As a part of the European Union's strategy for the sustainable use of plant protection products, the "IPM Decisions" project developed an online platform that gives farmers and advisors access to a wide range of DSS for major pests, weeds, and diseases in a variety of crops across Europe. Multiple DSS models relevant for different crops and geographical regions of Europe were selected for integration in the platform. Information on the models is compiled into a model catalogue, which serves as a core component of the IPM Decisions platform. To facilitate the use of these models, two application programming interfaces (APIs) were developed. In line with the FAIR (findable, accessible, interoperable, and reusable) principles, the DSS API provides access to models and their metadata, including descriptions of input and output parameters. The weather API enables access to European online weather data sources and adapts this data to meet the requirements of DSS models. While these APIs are integrated into the IPM decisions platform, they are also open source, allowing other crop protection and farm management software to inspect, download, modify, install, run, and use them. In this article, we describe the development of the DSS and weather APIs, outline their structure and definitions, and present the services that DSS API and weather API provide. Finally, we demonstrate their application through three practical use cases.
Net nitrogen (N) uptake results from active and water-mediated N inflows, partly offset by N exudation. These processes exhibit significant variations along root axes, suggesting the existence of specific zones of higher N exchanges. However, the precise location, drivers, and significance of such active root zones in root N budget remain unclear. Here, we identified and characterized active root zones using Root-CyNAPS (root-cycling nitrogen across plant scales), a new functional-structural plant model that simulates the space-time variations in net N uptake, based on the interactions between nitrogen, carbon, and water flows, and root anatomy in each segment of a 3D root system architecture. Our simulations on wheat for various plant ages and external nitrate concentrations revealed two zones of preferential net N uptake: one near the root apices and the other one coinciding with the lateral root initiation zone, both characterized by a net N uptake activity about 10 times higher than adjacent segments. Higher water uptake and N exudation rates were also located next to apices. The contribution of most active root zones (defined here by the top 10% of root length) varied between 20% and 80% of plant net N uptake, depending on root environment and age. Our simulations also showed that water-mediated N uptake could represent up to two-thirds of gross N uptake, while N losses could offset more than half of it. Root-CyNAPS offers coupling opportunities with other functional-structural models to simulate nitrogen, carbon, and water multiscale interactions across the soil-plant-atmosphere continuum.
Ensembles of multiple genomic prediction models have demonstrated improved prediction performance over the individual models contributing to the ensemble. The outperformance of ensemble models is expected from the Diversity Prediction Theorem, which states that for ensembles constructed with diverse prediction models, the ensemble prediction error becomes lower than the mean prediction error of the individual models. While a na & iuml;ve ensemble-average model provides baseline performance improvement by aggregating all individual prediction models with equal weights, optimizing weights for each individual model could further enhance ensemble prediction performance. The weights can be optimized based on their level of informativeness regarding prediction error and diversity. Here, we evaluated weighted ensemble-average models with three possible weight optimization approaches (linear transformation, Nelder-Mead and Bayesian) using flowering time and tillering traits from two maize nested associated mapping (NAM) datasets: TeoNAM and MaizeNAM. The three proposed weighted ensemble-average approaches improved prediction performance in several of the prediction scenarios investigated. In particular, the weighted ensemble models enhanced prediction performance when the adjusted weights differed substantially from the equal weights used by the na & iuml;ve ensemble models. For performance comparisons among the weighted ensembles, there was no clear superiority among the proposed approaches in both prediction accuracy and error across the prediction scenarios. Weight optimization for ensembles warrants further investigation to explore the opportunities to improve their prediction performance; for example, integration of a weighted ensemble with a simultaneous hyperparameter tuning process may offer a promising direction for further research.
Current crop growth models, whether process-based or data-driven, rarely incorporate spectral light composition, limiting their applicability in highly controlled environments such as vertical farming. This work enhances the predictive performance of a well-established process-based lettuce growth model by exploiting existing experimental evidence on the role of light spectrum in plant development. To this end, we introduce a parameter gamma, modelled as a function of key spectral features (i.e. the Blue:Red and Far-Red:Red ratios), selected through machine-learning techniques. The resulting adjusted model (aVH opt) is then validated on an independent literature dataset, showing a substantial reduction in prediction error compared to the reference model, with a more than 60% decrease in RMSE. The application of the aVH opt model to a commercial dataset confirms its capability to capture key spectral effects, but also reveals its sensitivity to environmental and biological variability not fully accounted for in the current formulation.
Root hairs are considered to be crucial for facilitating root water uptake, but the specific conditions influencing their impact remain unclear. We present a model that explores the efficacy of root hairs as a function of soil properties, root type, and tissue maturation. In particular, we investigate how the location of root hairs along the root axis, relative to root development, influences root water uptake. The analysis focuses on spatial distribution instead of individual traits such as hair density and length. Using a hydraulic model based on the finite-difference approach, water flow in single maize roots (lateral, seminal, crown, and brace) in two contrasting soil textures (sand and loam) was simulated. Soil water potential, root hair length, and the longitudinal extent of the root hair zone were varied to assess root hair impact on water uptake. (i) Root hairs become increasingly relevant in facilitating root water uptake as soil water potentials decrease. (ii) Root hair effectiveness is greater in sand than loam, due to the lower hydraulic conductivity of sand in the dry range. (iii) In contrast to loam, all root types increase total root water uptake with increasing rhizosphere conductance in sand drier than-2300 hPa. The positive role of hairs in water uptake depends on the profile of root conductivities along the root axis as long as the soil is more conductive than the root. Additionally, root hair functionality depends on soil hydraulic properties and hence is soil texture specific.
Rhizosphere models are typically 1D radial symmetric with the root as a boundary condition and root hairs as a reaction term (source/sink term). This rhizosphere representation comprises a model reduction for solute transport from 3D to 1D, which upscales the geometry of the root hairs. Further reduction is common in root architectural models, representing the root as point-sinks in a numerical mesh. We validated whether upscaling root hair geometry results in model-errors for nutrient uptake. We compared three levels of rhizosphere model reduction with varying geometric representation: (1) a 3D model with diffusion around root hairs, solved with a finite element method; (2) a 1D radial model with root hairs as sink term; (3) the root as point-sink. We investigated four different cases: short-thick-coarse, long-thick-coarse, short-thin-dense, and long-thin-dense root hairs. Each case was investigated for ammonium and phosphorus. In most cases, 1D and 3D solutions are close together, indicating that the 1D model reduction is valid. The 1D simulations deviate from 3D for slowly diffusing phosphorus taken up by long, thick, coarse root hairs; however, this is an extreme scenario and shows a theoretical limit. Further reducing the root to a point sink only estimates uptake for the ammonium scenario, which has fast diffusion. The root point sink is sensitive to the mesh resolution. Representing root hairs or roots by sink terms works for small radii or fast diffusing solutes. The 1D radial models can give valid and computationally fast results, but further reduction can lead to errors.
Process-based crop modelling platforms such as DSSAT are potentially valuable tools for crop breeding programmes, with the capacity to predict genotype-by-environment-by-management interactions. However, their application for breeding is challenged by the need to calibrate large numbers of genotypes within populations. In wheat (Triticum aestivum L.), using pre-existing DSSAT-CERES wheat ecotypes can introduce unrealistic parameter compensation during cultivar calibration. To address this, we developed a two-phase sequential calibration framework. This workflow uses phenotypic clustering to first define representative ecotypes using experiment-specific data before proceeding with cultivar-level parameter estimation. We demonstrate the utility of this framework to integrate direct measurements from proximal and remote sensing data collected on 14 genotypes grown under well-watered, drought, and heat stress field conditions. Incorporating experiment-derived ecotypes reduced compensatory adjustments in cultivar coefficients and improved simulation accuracy compared with default or non-representative ecotypes. Time-series data enhanced calibration, although the effect of different data combinations varied with environmental scenario and trait. Model simulations under stress conditions generally captured drought effects on biomass but underestimated heat stress impacts. This framework provides a systematic and scalable approach for integrating high-throughput phenotyping and process-based crop modelling.
Dwarf tomatoes are well suited for vertical farming due to their compact architecture and determinate growth, but their current genotypes are not optimized for high-density indoor systems. We developed the first functional-structural plant (FSP) model tailored to dwarf tomatoes in vertical farming to simulate light interception, photosynthesis, and biomass allocation at the organ level. The model integrates both static and dynamic modes using multiscale tree graph (MTG) formalism to encode plant architecture and was implemented in GroIMP. The model was parameterized and validated using experimental data obtained at multiple planting densities under controlled environmental conditions. The model accurately predicted key plant traits such as fruit dry mass, total dry mass, and leaf area index. In follow-up scenario analyses, we explored model responses to temperature perturbations (+/- 2 degrees C) and photosynthetically active radiation changes (+/- 20%) at selected planting densities, quantifying how these factors modulate the same traits. The model was executed both in dynamic and static mode to assess architectural ideotypes with focus on leaflet morphology, revealing that leaflet curvature and shape influence light distribution and carbon gain, with effects differing across densities and simulation modes. The model highlights the importance of dynamic feedbacks in high-density vertical farming and supports ideotype design through in silico evaluation of morphological traits. This work establishes a validated modelling framework for guiding breeding and cultivation strategies aimed at enhancing productivity and light use efficiency in vertical farming systems.
Accurate selection of favourable crop genotypes has motivated the exploration of diverse prediction algorithms for crop breeding applications. One genomic prediction method that has not been fully explored is graph attention networks (GAT). By directly analysing graphical data with the attention mechanism, GAT can incorporate the genotype-to-phenotype (G2P) structure to regularise predictions. As one potential G2P structure, a gene network can be inferred from interpretable machine learning models to effectively learn key features of prediction patterns, potentially improving prediction performance. Here, we investigated whether incorporating such data-driven prior knowledge into GAT improved prediction performance compared to GAT models representing a continuum of G2P structures, ranging from infinitesimal to fully connected. Applying the Diversity Prediction Theorem, we also combined these diverse G2P structures into an ensemble of GAT genomic prediction models to integrate complementary strengths of multiple models. The results for flowering time traits in two maize nested association mapping datasets showed a lack of consistent performance improvement in the data-driven prior knowledge GAT model. However, consistent outperformance was observed for the ensemble of GAT models. Improved predictions from the ensemble model may be driven by its ability to capture a more complete representation of the inferred gene network through the integration of information from diverse G2P structures. The observed results using the GAT methodology provided the foundation for potential performance improvement using GAT by integrating biological prior knowledge derived from omics data and empirically verified gene interactions in future research, thereby potentially enhancing the GAT ensemble performance.
This paper explores the application of data-driven system identification techniques in the frequency domain to obtain simplified, control-oriented models of photosynthesis regulation under oscillating light conditions. In-silico datasets are generated using simulations of the physics-based Basic DREAM Model (BDM) Funete et al.[2024], with light intensity signals – comprising DC (static) and AC (modulated) components as input and chlorophyll fluorescence (ChlF) as output. Using these data, the Best Linear Approximation (BLA) method is employed to estimate second-order linear time-invariant (LTI) transfer function models across different operating conditions defined by DC levels and modulation frequencies of light intensity. Building on these local models, a Linear Parameter-Varying (LPV) representation is constructed, in which the scheduling parameter is defined by the DC values of the light intensity, providing a compact state-space representation of the system dynamics.
Pest damage exhibits considerable spatial heterogeneity among individual plants in the field. While such spatial heterogeneity has often been treated as a nuisance in crop breeding trials, underlying biotic factors and loci remain poorly understood. To quantify the spatial variation in disease infection and associate it with neighboring genotypes, we applied two methods, Spatial Analysis of Field Trials with Splines (SpATS) and Neighbor Genome-Wide Association Study (Neighbor GWAS), to barley cultivars. Having compiled the CIMMYT Australia ICARDA Germplasm Evaluation (CAIGE) data, we first applied SpATS to three disease phenotypes such as the net form net blotch, spot form net blotch, and scald damage. This SpATS analysis showed extraneous phenotypic variation unexplained by smooth spatial trends, thereby leading us to focus on neighboring genotypes as an extraneous biological factor. We then applied the Neighbor GWAS model and found that neighbor genotypic identity explained 0.1–0.3 fractional variation in the three disease phenotypes. The Neighbor GWAS method also detected two significant or marginally significant variants on the barley 7H chromosome, which were associated with neighbor genotypic influence on the net form net blotch and scald damage. These variants were estimated to have beneficial effects that reduce disease damage by their allelic mixtures. Our findings suggest that neighbor genotypic identity can account for spatial variation in disease infection, providing a key to reduce pest damage by variety mixtures in field crops.
Crop yields in intercropping systems are the result of a combination of factors dominated by plastic responses in plant traits due to the heterogeneity associated with the intercrop design and row configuration. Disentangling their relative influence is infeasible in situ but crucial for cultivar selection and intercrop design. Using functional-structural plant (FSP) modelling, these effects can be separated in silico. Here, a mechanistic FSP model was developed, including three-dimensional aboveground plant architecture of maize and soybean, radiation distribution, and assimilate allocation using published data. The model was used to explore the potential to improve yields in a simultaneous intercrop by disentangling the contribution of three plastic traits related to photosynthesis, leaf thickness and plant height. The improved phenotypes were then simulated in two intercrop configurations for potential increases in land-use efficiency. The study revealed that for maize, photosynthesis had the greatest contribution (+78%), followed by plant height (+31%) and leaf thickness (+6%), where the total maize monoculture phenotype produced the greatest maize yield without affecting the yield of intercropped soybean. However, soybean trait plasticity had a negligible effect on soybean productivity, but its monoculture phenotype with a low light-saturated photosynthetic rate resulted in the greatest intercropped maize yield. These improved phenotypes may increase land-use efficiency by 1%-3% relative to the standard monoculture systems of the Midwest, USA, which is also a similar to 22% increase from published empirical data. Together, these results could aid the selection of crop germplasm to improve the productivity of a simultaneous intercrop.
PlantCoreMetabolism is a manually curated charge- and proton-balanced model of plant primary metabolism published in 2018. Since it consists of reactions common to all plants, the model has been used to study metabolism in a wide range of plant systems, including C3 leaves, crassulacean acid metabolism leaves, fruit pericarp cells, maize roots, and stomatal guard cells. Here, we summarize the application of diel flux balance analysis (FBA) leaf models derived from the PlantCoreMetabolism model. Diel FBA models extend traditional FBA by optimizing metabolism over a 24-h light-dark cycle, capturing metabolic processes spread across multiple temporal phases. In this methods paper, we also describe key use cases of diel FBA leaf models, provide case study tutorials to demonstrate the utility of a diel leaf model, and introduce flux visualization tools to facilitate future work with the model.