Abstract. Australia spans nearly the full spectrum of global bioclimatic zones, from tropical savannas to arid deserts and alpine environments. Understanding how climate constrains vegetation growth across this gradient is essential for interpreting ecosystem dynamics and informing land-management decisions. We introduce a Plant Growth Index (PGI), a continental-scale metric derived from meteorological data from the Bureau of Meteorology Atmospheric high-resolution Regional Reanalysis for Australia (BARRA-R), European Space Agency (ESA) Plant Functional Type (PFT) layers, and C3/C4 grass fractions estimated from NASA Moderate Resolution Imaging Spectroradiometer (MODIS) enhanced vegetation index (EVI). The Plant Growth Index captures year-to-year variation in vegetation status, water availability, and climatic conditions. Spatially, values of the PGI are highest in tropical and subtropical regions and lowest in arid deserts. Benchmarking against gross primary productivity (GPP) from the Terrestrial Ecosystem Research Network (TERN) OzFlux network, the Normalised Difference Vegetation Index (NDVI), the Standardised Precipitation-Evapotranspiration Index (SPEI), and Australian Grassland and Rangeland Assessment by Spatial Simulation (Aussie-GRASS) indicates that PGI broadly reflects regional vegetation productivity patterns. The PGI provides a reproducible, continental-scale tool for ecological modelling, rangeland monitoring, and climate-impact studies and can be accessed at https://doi.org/10.5281/zenodo.18762343 (Retkute et al., 2026).
Reliable inference in infectious disease modeling requires careful treatment of both model structure and the relationship between latent infection dynamics and observed data. Likelihood functions, which link model parameters to empirical observations, can be formulated either to explicitly represent underlying disease transmission and reporting processes (process-based) or to summarize statistical patterns in aggregated outcomes (observation-based). Stochastic models capture inherent variability in transmission and detection, whereas deterministic models describe average system behavior and often rely on statistical assumptions to account for residual uncertainty. Using two neglected tropical disease (NTD) models, we compare parameter estimation based on complete individual-level events with that based on aggregated counts. By generating synthetic outbreak data from stochastic simulations and analyzing it under alternative modeling frameworks, we show how different combinations of model formulation and likelihood structure influence both point estimates and uncertainty quantification. Our findings indicate that, even when detailed process information is unavailable, observation-based likelihoods can produce robust parameter estimates and credible uncertainty intervals, highlighting their usefulness for practical decision-making in contexts with limited or aggregated surveillance data.
Accurate parameter estimation is fundamental to quantitative epidemiology as it provides the foundation for robust modelling and evidence-based decision-making. We present a novel two-stage framework, that is designed to enhance computational efficiency for parameter estimation of complex, stochastic, or spatially explicit models. Traditional Approximate Bayesian Computation (ABC) methods often face prohibitive costs when likelihoods are analytically intractable and acceptance rates are low. Our hybrid approach, ABC-RF-rejection, integrates ABC rejection sampling with Random Forest (RF) classification to selectively identify parameter sets likely to satisfy observed data constraints. In the first stage, a small-scale ABC rejection step generates a labelled training dataset of accepted and rejected particles. In the second stage, a trained RF model decouples posterior exploration from expensive forward simulations by predicting acceptance probabilities for a substantially larger candidate set. This allows the algorithm to focus computational resources on high-probability regions of the parameter space. We apply the ABC-RF-rejection approach to three distinct epidemiological contexts: stochastic simulations of onchocerciasis vector control, spatially explicit modelling of cassava brown streak virus spread in Uganda, and the 2014-2015 West Africa Ebola outbreak in heterogeneous populations. This framework provides an adaptable and robust solution for rapid, evidence-based decision-making in settings characterised by spatial heterogeneity, stochasticity, and limited surveillance data.
Cassava is a critical staple crop for food security and rural livelihoods in Sub-Saharan Africa, yet high-resolution maps of its distribution remain scarce, particularly for smallholder systems. In this study, we generated a 10 m resolution cassava presence map for Uganda (CM24) by fine-tuning a Random Forest classifier on TESSERA foundation model embeddings derived from Sentinel-1 and Sentinel-2 time series. Using field survey data from the Copernicus4GEOGLAM campaign for training and validation, the model achieved excellent discriminative ability (validation AUC = 0.9532, test AUC = 0.9524). Visual validation against high-resolution satellite imagery confirmed good spatial agreement, capturing both large contiguous fields and small fragmented plots. Comparison with two existing global products (CassavaMap and SPAM2020) and two seasons of national survey data conducted by the Uganda Bureau of Statistics showed that CM24 produced national harvested area estimates that fell between the two survey totals, whereas CassavaMap and SPAM2020 systematically overestimated harvested area by factors of two to three. Our results demonstrate that foundation-model embeddings offer a robust and scalable approach for mapping cassava in heterogeneous smallholder landscapes. The resulting CM24 map provides a spatially explicit tool to support disease surveillance, agricultural monitoring, and food security planning in Uganda and beyond.
Abstract Banana bunchy top virus (BBTV) continues to threaten smallholder livelihoods and food security across sub-Saharan Africa. While clean-seed programmes are widely promoted, their long-term effectiveness is often compromised by rapid reinfection in endemic landscapes. We developed an integrated framework combining spatially explicit, stochastic epidemiological modelling with additional cost–benefit analysis and use of socio-behavioural data to evaluate strategies for stabilising production. Our results demonstrate that frequent monthly inspections and accurate symptom detection are essential for disease suppression. Crucially, the economic analysis reveals that prioritizing diagnostic competence as an economic asset is necessary: improving detection efficiency can more than double farmer net revenue under realistic market conditions. Socio-behavioural findings further confirm that a farmer’s ability to correctly recognise symptoms is the strongest predictor of roguing adoption, far outweighing demographic characteristics. These results provide quantifiable guidance for disease management and highlight that the sustainability of clean-seed interventions hinges on shifting policy from simple seed replacement to investing in farmer diagnostic capacity. Strengthening this local surveillance capability transforms informal seed systems into resilient, durable tools for safeguarding household nutrition and regional food security.
Bananas and plantains are among the world’s most important staple food crops and provide daily calories, income, and nutritional security for millions of smallholder households, particularly across sub-Saharan Africa. Accurately estimating epidemiological parameters for major threats such as banana bunchy top virus (BBTV) is essential for predicting disease spread and designing effective management strategies, yet the limited and resource-constrained field data available in smallholder systems make this extremely challenging. Here, we introduce a data-augmented Adaptive Multiple Importance Sampling (DA-AMIS) framework that integrates Bayesian inference with a mechanistic epidemic model to recover key BBTV transmission parameters from small field experiments. Using detailed individual-level observations from a 24-plant experiment in Benin, we jointly infer latent infection times, aphid-mediated dispersal characteristics, and primary and secondary transmission rates. We validate these estimates against independent BBTV datasets from Burundi and Malawi, finding close correspondence between simulated and observed prevalence trajectories, demonstrating the transferability of inferred parameters across regions. Our results indicate that approximately 12% of replanting suckers are infected at planting, emphasizing the high risk of BBTV introduction through planting material, and simulations identify April as the period of peak infection pressure, providing actionable insight for surveillance timing. These findings show that small field experiments, when combined with advanced Bayesian computational methods, can yield robust and generalizable epidemiological parameter estimates.
Banana bunchy top disease (BBTD), caused by banana bunchy top virus (BBTV), threatens food security and livelihoods across sub-Saharan Africa. Effective surveillance is critical for early detection, but in many regions, monitoring is constrained to short operational windows due to limited funding and laboratory capacity. When surveillance is restricted to a single calendar year, sampling effort must be carefully allocated to maximize early detection, yet quantitative guidance linking detection performance to cost is lacking. To address this, we simulated the spatio-temporal spread of BBTV in Benin and retrospectively evaluated country-level BBTD surveillance using a single-year cross-sectional survey design. We found that early detection is feasible but prohibitively expensive (USD 100,000 per year). Under a constrained budget of USD 10,000 per year, an optimal strategy - defined as reaching 75% mean detection probability earliest - was 500 sites with 10 samples per site, though this could delay BBTD detection by up to one year. Our integrated simulation - economic framework quantifies trade-offs between cost and detection likelihood, providing guidance for resource-efficient national-scale BBTD surveillance and a transferable approach for plant pathogen monitoring in smallholder systems. ### Competing Interest Statement The authors have declared no competing interest. Gates Foundation, https://ror.org/0456r8d26, INV070408
Wheat stem rust is a fungal disease that can cause total loss of a wheat crop. Outbreaks are typically far-reaching because infectious spores can disperse aerially over thousands of kilometres. Rapid response to an outbreak requires a timely and accurate estimate of disease risk. We constructed an integrated mechanistic model of stem rust infection driven by seven-day meteorological forecasts and near-real time disease surveillance to inform mitigation at the national scale. Focusing on Ethiopia as the largest wheat producer of sub-Saharan Africa, yet where food insecurity is high, we evaluate the predictability of wheat stem rust infections during the rain-fed seasons of 2015–2022. The epidemiological model performs well at predicting disease observations two weeks ahead of time, from cross-validation root-mean-square-error values were in the range 0.25–0.35, ROC skill score 0.58–0.68, and odds ratio skill score 0.78–0.90. Results of alternative model scenarios show the benefit of assimilating recent surveys and simulating spore dispersal to achieve a good prediction. When driven by seven-day weather forecasts, the model therefore provides up to 21 days to respond to likely stem rust infections and prevent epidemic spread and subsequent crop loss. The mechanistic form and large scale of the epidemiological model has utility beyond in-season predictions, particularly with respect to scenario testing for practical interventions. The results demonstrated that long-distance dispersal has an important role in the spread of stem rust in a region of diverse agro-ecology. The forecast system presented here can be adapted to other plant pathosystems where aerial dispersal is the primary risk of disease spread.
Banana diseases impose severe production losses in tropical smallholder farming systems, yet accurate in-field visual diagnosis remains difficult: symptom expression varies across cultivars and growth stages, and several diseases produce morphologically overlapping foliar signs. We developed a probabilistic image-recognition framework for detecting five economically important banana diseases - Xanthomonas Wilt, Banana Bunchy Top Disease, Fusarium Wilt (Panama disease), Yellow Sigatoka, and Black Sigatoka - from in-field photographs, without any disease-specific fine-tuning of the vision backbone. The approach extracts frozen 1,152-dimensional embeddings from the DINOv3 vision foundation model and couples them with a conditional normalizing flow, trained on four publicly available datasets spanning diseased banana plants, healthy tissue, non-banana vegetation, and general natural imagery. On an independent test set the model achieved F1 scores exceeding 0.98, average precision values of 0.968-0.999, and AUROC values of 0.997-1.0 across all five diseases evaluated as binary detection problems. Multi-class accuracy was near-perfect, with limited confusion between Yellow Sigatoka and Black Sigatoka - a biologically plausible ambiguity attributable to overlapping early-infection foliar symptoms. Because the normalizing flow estimates explicit conditional probability densities rather than decision boundaries, two complementary log-likelihood ratios can be derived: a disease ratio comparing each disease class against healthy banana, and a plant ratio comparing banana against non-banana imagery. Together these define an interpretable two-dimensional diagnostic space that simultaneously quantifies evidence for disease presence and image relevance, cleanly separating diseased plants, healthy plants, and out-of-distribution images while flagging uncertain predictions for confirmatory testing. Inference on frozen embeddings is lightweight and compatible with smartphone deployment, providing a scalable, uncertainty-aware diagnostic tool for smallholder farming systems and disease surveillance programmes.
We introduce a novel two-stage parameter estimation framework designed to improve computational efficiency in settings involving complex, stochastic, or analytically intractable dynamic models. The proposed method, termed ABC-RF-rejection, integrates Approximate Bayesian Computation (ABC) rejection sampling with Random Forest (RF) classification to efficiently screen parameter sets that produce simulations consistent with observed data. We evaluate the performance of this approach using both a deterministic Susceptible-Infected-Removed (SIR) epidemic model and a spatially explicit stochastic epidemic model. Results indicate that ABC-RF-rejection achieves substantial gains in computational efficiency while maintaining parameter inference accuracy comparable with standard ABC rejection methods. Finally, we apply the algorithm to estimate parameters governing the spatial spread of cassava brown streak disease (CBSD) in Nakasongola district, Uganda.
Summary Societal Impact Statement Banana and plantain are critical for food security and income generation in West Africa, particularly for smallholder farmers in countries such as Benin. However, banana bunchy top virus (BBTV) poses a major threat to sustainable production. This study integrates high-resolution remote sensing, field epidemiology, mathematical modelling, and socioeconomic analysis to improve understanding of the spread of BBTV and to evaluate the effectiveness of different disease control strategies. Our findings reveal that BBTV risk is highest in southern Benin, where socioeconomic vulnerability is also greatest. Low disease awareness, limited adoption of effective control methods and informal exchange of planting material exacerbate this risk. By identifying priority areas and strategies tailored to local social, economic, and agroecological contexts, this research offers a roadmap for designing targeted, sustainable BBTV management programs. These insights can support smallholder resilience, reduce disease burden, and safeguard banana-based livelihoods across sub-Saharan Africa. ### Competing Interest Statement The authors have declared no competing interest. Bill & Melinda Gates Foundation, https://ror.org/0456r8d26, INV070408, INV010652
West Nile virus (WNV) is the most common cause of mosquito-borne disease in the continental USA, with an average of 1200 severe, neuroinvasive cases reported annually from 2005 to 2021 (range 386–2873). Despite this burden, efforts to forecast WNV disease to inform public health measures to reduce disease incidence have had limited success. Here, we analyze forecasts submitted to the 2022 WNV Forecasting Challenge, a follow-up to the 2020 WNV Forecasting Challenge. Forecasting teams submitted probabilistic forecasts of annual West Nile virus neuroinvasive disease (WNND) cases for each county in the continental USA for the 2022 WNV season. We assessed the skill of team-specific forecasts, baseline forecasts, and an ensemble created from team-specific forecasts. We then characterized the impact of model characteristics and county-specific contextual factors (e.g., population) on forecast skill. Ensemble forecasts for 2022 anticipated a season at or below median long-term WNND incidence for nearly all (> 99
Banana is an important cash and food crop worldwide. Recent outbreaks of banana diseases are threatening the global banana industry and smallholder livelihoods. Remote sensing data offer the potential to detect the presence of disease, but formal analysis is needed to compare inferred disease data with observed disease data. In this study, we present a novel remote-sensing-based framework that combines Landsat-8 imagery with meteorology-informed phenological models and machine learning to identify anomalies in banana crop health. Unlike prior studies, our approach integrates domain-specific crop phenology to enhance the specificity of anomaly detection. We used a pixel-level random forest (RF) model to predict 11 key vegetation indices (VIs) as a function of historical meteorological conditions, specifically daytime and nighttime temperature from MODIS and precipitation from NASA GES DISC. By training on periods of healthy crop growth, the RF model establishes expected VI values under disease-free conditions. Disease presence is then detected by quantifying the deviations between observed VIs from Landsat-8 imagery and these predicted healthy VI values. The model demonstrated robust predictive reliability in accounting for seasonal variations, with forecasting errors for all VIs remaining within 10% when applied to a disease-free control plantation. Applied to two documented outbreak cases, the results show strong spatial alignment between flagged anomalies and historical reports of banana bunchy top disease (BBTD) and Fusarium wilt Tropical Race 4 (TR4). Specifically, for BBTD in Australia, a strong correlation of 0.73 was observed between infection counts and the discrepancy between predicted and observed NDVI values at the pixel with the highest number of infections. Notably, VI declines preceded reported infection rises by approximately two months. For TR4 in Mozambique, the approach successfully tracked disease progression, revealing clear spatial spread patterns and correlations as high as 0.98 between VI anomalies and disease cases in some pixels. These findings support the potential of our method as a scalable early warning system for banana disease detection.
Epidemics of Banana Bunchy Top Disease (BBTD) in sub-Saharan Africa are threatening global food security and endangering the livelihoods of smallholder farmers. This study introduces methods for developing data-based models to derive banana production maps and process-based models to assess the potential spread of BBTV at a landscape scale. We introduce two novel aspects: a methodology for deriving probabilistic banana production maps based on high-resolution remote sensing products and parameterization of the epidemiological model for BBTD from limited survey data. We generated a countrywide banana production map for Tanzania and a state-wide map for Ogun State in Nigeria. We used the banana map together with published data from BBTD surveys to parameterize a model for BBTD spread in Tanzania. Our results emphasize the importance of surveys, as having data on the presence and absence of Banana Bunchy Top Virus (BBTV) at different stages of epidemics is crucial not only for effective control of the disease but also for prediction, including making reasonable model assumptions, model parameterization, and model validation that underpin predictions.
Accurate estimation of epidemiological parameters from limited field data remains a major challenge in plant disease modeling. We present a novel data-augmented adaptive multiple importance sampling (DA-AMIS) framework that integrates Bayesian inference with stochastic epidemic modeling to estimate key transmission parameters from small field experiments. Using detailed individual-level observations from a 24-plant experiment on the natural spread of banana bunchy top virus (BBTV) in Benin, we jointly inferred infection timing, dispersal characteristics, and transmission rates for both primary and secondary infections. Model validation against independent datasets from BBTV field trials in Burundi and Malawi showed close correspondence between simulated and observed prevalence dynamics, confirming the generality of parameter estimates across regions. The inferred 12% infection rate of replanting suckers underscores the risk of disease introduction through planting material, while simulations identified April as the period of peak infection, providing actionable insights for surveillance timing. ### Competing Interest Statement The authors have declared no competing interest. Gates Foundation, https://ror.org/0456r8d26, INV070408, INV010652 The Mastercard Foundation and University of Cambridge Climate Resilience and Sustainability Research Fund
Estimating parameters for multiple datasets can be time consuming, especially when the number of datasets is large. One solution is to sample from multiple datasets simultaneously using Bayesian methods such as adaptive multiple importance sampling (AMIS). Here, we use the AMIS approach to fit a von Mises distribution to multiple datasets for wind trajectories derived from a Lagrangian Particle Dispersion Model driven from 3D meteorological data. A posterior distribution of parameters can help to characterise the uncertainties in wind trajectories in a form that can be used as inputs for predictive models of wind-dispersed insect pests and the pathogens of agricultural crops for use in evaluating risk and in planning mitigation actions. The novelty of our study is in testing the performance of the method on a very large number of datasets (>11,000). Our results show that AMIS can significantly improve the efficiency of parameter inference for multiple datasets.
There is an urgent need for mathematical models that can be used to inform the deployment of surveillance, early warning and management systems for transboundary pest invasions. This is especially important for desert locust, one of the most dangerous migratory pests for smallholder farmers. During periods of desert locust upsurges and plagues, gregarious adult locusts form into swarms that are capable of long-range dispersal. Here we introduce a novel integrated modelling framework for use in predicting gregarious locust populations. The framework integrates the selection of breeding sites, maturation through egg, hopper and adult stages and swarm dispersal in search of areas suitable for feeding and breeding. Using a combination of concepts from epidemiological modelling, weather and environment data, together with an atmospheric transport model for swarm movement we provide a tool to forecast short- and long-term swarm movements. A principal aim of the framework is to provide a practical starting point for use in the next upsurge.
Acclimation of photosynthesis to light intensity (photoacclimation) takes days to achieve and so naturally fluctuating light presents a potential challenge where leaves may be exposed to light conditions that are beyond their window of acclimation. Experiments generally have focused on unchanging light with a relatively fixed combination of photosynthetic attributes to confer higher efficiency in those conditions. Here a controlled LED experiment and mathematical modelling was used to assess the acclimation potential of contrasting Arabidopsis thaliana genotypes following transfer to a controlled fluctuating light environment, designed to present frequencies and amplitudes more relevant to natural conditions. We hypothesize that acclimation of light harvesting, photosynthetic capacity and dark respiration are controlled independently. Two different ecotypes were selected, Wassilewskija-4 (Ws), Landsberg erecta (Ler) and a GPT2 knock out mutant on the Ws background (gpt2-), based on their differing abilities to undergo dynamic acclimation i.e. at the sub-cellular or chloroplastic scale. Results from gas exchange and chlorophyll content indicate that plants can independently regulate different components that could optimize photosynthesis in both high and low light; targeting light harvesting in low light and photosynthetic capacity in high light. Empirical modelling indicates that the pattern of 'entrainment' of photosynthetic capacity by past light history is genotype-specific. These data show flexibility of photoacclimation and variation useful for plant improvement.
Objective Applications of machine learning in healthcare are of high interest and have the potential to improve patient care. Yet, the real-world accuracy of these models in clinical practice and on different patient subpopulations remains unclear. To address these important questions, we hosted a community challenge to evaluate methods that predict healthcare outcomes. We focused on the prediction of all-cause mortality as the community challenge question. Materials and methods Using a Model-to-Data framework, 345 registered participants, coalescing into 25 independent teams, spread over 3 continents and 10 countries, generated 25 accurate models all trained on a dataset of over 1.1 million patients and evaluated on patients prospectively collected over a 1-year observation of a large health system. Results The top performing team achieved a final area under the receiver operator curve of 0.947 (95% CI, 0.942-0.951) and an area under the precision-recall curve of 0.487 (95% CI, 0.458-0.499) on a prospectively collected patient cohort. Discussion Post hoc analysis after the challenge revealed that models differ in accuracy on subpopulations, delineated by race or gender, even when they are trained on the same data. Conclusion This is the largest community challenge focused on the evaluation of state-of-the-art machine learning methods in a healthcare system performed to date, revealing both opportunities and pitfalls of clinical AI.