Satellite-based land cover maps fail to distinguish between natural forests and planted trees when certifying agricultural plantations as deforestation-free, a critical need for smallholder farmers facing emerging legislation around sustainable supply chains (such as the European Union Deforestation Regulation). Previous efforts to map tree plantations exclude smallholder farms, which require intra-field detail about individual tree crowns that can only be observed in costly, very-high resolution (VHR, <1m) satellite imagery, and limit scalability. To provide a replicable benchmark for mapping tree crop plantations, we introduce a novel earth observations dataset of hierarchically-labeled plantation and forest locations across the African continent. We compile for these points a unique feature set combining a VHR satellite-derived canopy height product and open-source multispectral and synthetic aperture radar data, to fuse high spatial resolution with multitemporal, multiscale, and multimodal features. Our dataset has continental coverage for Africa's plantation-growing nations across a diverse range of climate zones, with more delineation for species and geographical region than any other related tree dataset. Samples are structured with different domains to account for real-world distribution shifts to study the interplay of domain generalization with various feature attributes of remotely sensed data. We test how pre-trained models can perform on the task of tree plantation mapping, and benchmark performance at scale and across different distributions.
Air pollution from coal electricity generation is a major driver of poor air quality in India and its effects on human health have been extensively studied. Despite considerable evidence that the same pollution also reduces crop productivity, we lack similar quantitative assessments of coal electricity's crop damages. Here, we estimate rice and wheat crop losses from coal generation's nitrogen dioxide (NO2) emissions using a regression model that combines station-level electricity generation and wind direction, satellite-measured NO2, and its association with crop productivity. Coal emissions impact yields up to 100 km away from power stations. In parts of West Bengal, Madhya Pradesh, and Uttar Pradesh heavily exposed to coal-linked NO2, annual yield losses exceed 10%, equivalent to approximately 6 y worth of average annual yield growth in both rice and wheat in India between 2011 and 2020. While station-specific crop damages (value of lost output) are almost always lower than mortality damages (monetized value of annual premature PM2.5-related deaths), crop damage intensity (crop damage per GWh of electricity generated) is frequently higher than mortality damage intensity (mortality damage/GWh). Rice damage intensity exceeds mortality damage intensity at 58, and wheat damage intensity at 35 of the 144 power stations studied. The stations associated with the largest crop losses differ from those associated with the highest mortality. Co-optimizing for crop gains and mortality reduction slightly increases and meaningfully changes the distribution of social benefits from reducing emissions, highlighting the importance of considering crop losses alongside health impacts when regulating coal electricity emissions in India.
Farmers in the USA have rapidly expanded the use of cover crops, with the national cover crop area nearly doubling since 2012. Despite many benefits that motivate public subsidies, questions remain about potential downsides. Here, using satellite observations from over 100,000 fields, half of which recently adopted cover crops, we demonstrate both positive and negative impacts of cover cropping, including: (1) declines in average yields for corn and soybean, by similar to 3% and similar to 2%, respectively; (2) delays in planting of corn (4 days) and soybean (2.5 days); and (3) reduced damages in the wet spring of 2019, with cover crop fields only half as likely to experience prevented planting as non-cover-crop fields. Cover cropping appears to reduce important aspects of farmer risk in wet conditions but increase them in dry conditions. Timely planting of the cash crop deserves emphasis moving forward, as we show eliminating planting delays would reduce yield penalties by roughly 50% for corn and 90% for soybean.
Parameter-efficient fine-tuning (PEFT) techniques such as low-rank adaptation (LoRA) can effectively adapt large pre-trained foundation models to downstream tasks using only a small fraction (0.1%-10%) of the original trainable weights. An under-explored question of PEFT is in extending the pre-training phase without supervised labels; that is, can we adapt a pre-trained foundation model to a new domain via efficient self-supervised pre-training on this domain? In this work, we introduce ExPLoRA, a highly effective technique to improve transfer learning of pre-trained vision transformers (ViTs) under domain shifts. Initializing a ViT with pre-trained weights on large, natural-image datasets such as from DinoV2 or MAE, ExPLoRA continues the unsupervised pre-training objective on a new domain, unfreezing 1-2 pre-trained ViT blocks and tuning all other layers with LoRA. We then fine-tune the resulting model only with LoRA on this new domain for supervised learning. Our experiments demonstrate state-of-the-art results on satellite imagery, even outperforming fully pre-training and fine-tuning ViTs. Using the DinoV2 training objective, we demonstrate up to 8% improvement in linear probing top-1 accuracy on downstream tasks while using <10% of the number of parameters that are used in prior fully-tuned state-of-the art approaches. Our ablation studies confirm the efficacy of our approach over other baselines such as PEFT. Code is available at: https://samar-khanna.github.io/ExPLoRA/
Building sustainable food systems that are resilient to climate change will require improved agricultural management and policy. One common practice that is well-known to benefit crop yields is crop rotation, yet there remains limited understanding of how the benefits of crop rotation vary for different crop sequences and for different weather conditions. To address these gaps, we leverage crop type maps, satellite data, and causal machine learning to study how precrop effects on subsequent yields vary with cropping sequence choice and weather. Complementing and going beyond what is known from randomized field trials, we find that (i) for those farmers who do rotate, the most common precrop choices tend to be among the most beneficial, (ii) the effects of switching from a simple rotation (which alternates between two crops) to a more diverse rotation were typically small and sometimes even negative, (iii) precrop effects tended to be greater under rainier conditions, (iv) precrop effects were greater under warmer conditions for soybean yields but not for other crops, and (v) legume precrops conferred smaller benefits under warmer conditions. Our results and the methods we use can enable farmers and policy makers to identify which rotations will be most effective at improving crop yields in a changing climate.
Low fertilizer use by smallholder farmers continues to limit crop yields in sub-Saharan Africa. We study the longer-term outcomes of a field experiment conducted between 2014 and 2016, which found that plot-specific fertilizer recommendations combined with a subsidy increased fertilizer use and maize yields. We return in 2019 and find that effects dissipate after the subsidy is discontinued. Our follow-up results suggest that credit constraints strongly limit investment in fertilizer, because farmers have received information about what fertilizer types and amounts to apply and have an experience of fertilizer as profitable. We find that the 2016 treatment effects were driven by the most productive farmers—those with more fertile soils who cultivated larger plots of land. Our analysis features use of both self-reported and satellite-derived yield estimates. Our results suggest the potential importance of sustained financial support in combination with information to induce smallholder farmers to continue to invest in fertilizers.
Efforts to anticipate and adapt to future climate can benefit from historical experiences. We examine agroclimatic conditions over the past 50 y for five major crops around the world. Most regions experienced rapid warming relative to interannual variability, with 45% of summer and 32% of winter crop area warming by more than two SD (σ). Vapor pressure deficit (VPD), a key driver of plant water stress, also increased in most temperate regions but not in the tropics. Precipitation trends, while important in some locations, were generally below 1σ. Historical climate model simulations show that observed changes in crops' climate would have been well predicted by models run with historical forcings, with two main surprises: i) models substantially overestimate the amount of warming and drying experienced by summer crops in North America, and ii) models underestimate the increase in VPD in most temperate cropping regions. Linking agroclimatic data to crop productivity, we estimate that climate trends have caused current global yields of wheat, maize, and barley to be 10, 4, and 13% lower than they would have otherwise been. These losses likely exceeded the benefits of CO2 increases over the same period, whereas CO2 benefits likely exceeded climate-related losses for soybean and rice. Aggregate global yield losses are very similar to what models would have predicted, with the two biases above largely offsetting each other. Climate model biases in reproducing VPD trends may partially explain the ineffectiveness of some adaptations predicted by modeling studies, such as farmer shifts to longer maturing varieties.
Efforts to combat land degradation globally have led to the widespread promotion of sustainable land management practices (SLMPs) aimed at reducing surface runoff and erosion. Despite their extensive implementation, long-term evaluations of these practices are limited, especially in data-scarce regions. In our study, we assess the long-term impact of large-scale SLMPs in Ethiopia using remotely sensed data from the past 24 years on 122 watersheds. Using a synthetic control method that does not require an explicit control group, we find statistically significant positive effects of SLMPs in both wet and dry seasons. These benefits persist at least eight years beyond the intervention period. Our findings highlight the need for multi-season impact assessments. Focusing only on the wet season may overlook key outcomes in dryland regions, underestimating the effectiveness of large-scale, multi-year projects. We further find that effects were most positive in drought-prone agricultural highlands, and that some administrative zones appear more effective than others at implementation. Efficient and affordable monitoring of sustainable agricultural water and land management and watershed conservation is crucial for understanding which interventions are effective and can provide opportunities for alternative financing mechanisms.
The preference for simple explanations, known as the parsimony principle, has long guided the development of scientific theories, hypotheses, and models. Yet recent years have seen a number of successes in employing highly complex models for scientific ...
Southern Africa faces high food insecurity and projected declines in agroclimatic conditions. Multiple satellite measures indicate that cropland productivity has stagnated for most of the region except South Africa in the past 20 years, in contrast to what official crop statistics suggest. Climate trends do not explain this stagnation, with the region experiencing more rainfall and less warming than most climate model projections. A change of course is needed before climate impacts accelerate.
The interdependence of climate change and agricultural land use remains a critical, yet unquantified, area of concern for future food production. Here we determine climate-driven cropland change based on an empirical model of cropland response to changes in agricultural productivity. By estimating counterfactual total factor productivity in a scenario without climate change, we find that 88 million hectares (90
Increasing agricultural productivity is a gradual process with significant time lags between research and development (R&D) investment and the resulting gains. We estimate the response of US agricultural Total Factor Productivity to both R&D investment and weather and quantify the public R&D spending required to offset the emerging impacts of climate change. We find that offsetting the climate-induced productivity slowdown by 2050 will require R&D spending over 2021 to 2050 to grow at 5.2 to 7.8% per year under a fixed spending growth scenario or by an additional $2.2 to $3.8B per year under a fixed supplement spending scenario (in addition to the current spending of ~$5B per year). This amounts to an additional $208 to $434B or $65 to $113B over the period, respectively, and would be comparable in ambition to the public R&D spending growth that followed the two World Wars.
Sugarcane is an important source of food, biofuel, and farmer income in many countries. At the same time, sugarcane is implicated in many social and environmental challenges, including water scarcity and nutrient pollution. Currently, few of the top sugar-producing countries generate reliable maps of where sugarcane is cultivated. To fill this gap, we introduce a dataset of detailed sugarcane maps for the top 13 producing countries in the world, comprising nearly 90 % of global production. Maps were generated for the 2019-2022 period by combining data from Global Ecosystem Dynamics Investigation (GEDI) and Sentinel-2 (S2). GEDI data were used to provide training data on where tall and short crops were growing each month, while S2 features were used to map tall crops for all cropland pixels each month. Sugarcane was then identified by leveraging the fact that, among all non-tree species grown in cropland areas, sugarcane is typically tall for the largest fraction of time. Comparisons with field data, pre-existing maps, and official government statistics all indicated high precision and high recall of our maps. Agreement with field data at the pixel level exceeded 80 % in most countries, and subnational sugarcane areas from our maps were consistent with government statistics. Exceptions appeared mainly due to problems in underlying cropland masks or due to under-reporting of sugarcane area by governments. The final maps should be useful in studying the various impacts of sugarcane cultivation and producing maps of related outcomes such as sugarcane yields.
Ongoing advances in satellite remote sensing data and machine learning methods have enabled crop yield estimation at various spatial and temporal resolutions. While yield mapping at broader scales (e.g., state or county level) has become common, mapping at finer scales (e.g., field or subfield) has been limited by the lack of ground truth data for model training and evaluation. Here we present a scale transfer framework, named Quantile loss Domain Adversarial Neural Networks (QDANN), that leverages knowledge from county-level datasets to map crop yields at the subfield level. Based on the strategy of unsupervised domain adaptation, QDANN is trained on labeled county-level data and unlabeled subfield-level data, with no requirement for yield information at the subfield level. We evaluate the proposed method applied to Landsat imagery and Gridmet weather data for maize, soybean, and winter wheat fields in the United States, using as reference data yield monitor records from roughly one million field-year observations. The model is compared with several process- based and machine learning-based benchmark approaches that train on simulated yield records or county-level data. QDANN-estimated yields achieved an R2 2 score (RMSE) of 48 % (2.29 t/ha), 32 % (0.85 t/ha), and 39 % (1.40 t/ha) for maize, soybean, and winter wheat in comparison with the ground-based yield measures, respectively. These performances are higher than benchmark approaches and are nearly as good as models trained on field-level data. When aggregated to the county level, the improvement achieved by QDANN is more pronounced and the R2 2 scores (RMSE) improved to 78 % (0.98 t/ha), 62 % (0.37 t/ha), and 53 % (1.00 t/ha) for maize, soybean, and winter wheat, respectively. This study demonstrates that the proposed scale transfer framework can serve as a reliable approach for yield mapping at the subfield level when there is no access to fine- scale yield information. Based on the QDANN model, we have generated and made publicly available 30-m annual yield maps for major crop-producing states in the U.S. since 2008.
Index insurance is a promising tool to reduce the risk faced by farmers, but high basis risk, which arises from imperfect correlation between the index and individual farm yields, has limited its adoption to date. Basis risk arises from two fundamental sources: the intrinsic heterogeneity within an insurance zone (zonal risk), and the lack of predictive accuracy of the index (design risk). Whereas previous work has focused almost exclusively on design risk, a theoretical and empirical understanding of the role of zonal risk is still lacking. Here we investigate the relative roles of zonal and design risk, using the case of maize yields in Kenya. Our first contribution is to derive a formal decomposition of basis risk, providing a simple upper bound on the insurable basis risk that any index can reach within a given zone. Our second contribution is to provide the first large-scale empirical analysis of the extent of zonal versus design risk. To do so, we use satellite estimates of yields at 10m resolution across Kenya, and investigate the effect of using smaller zones versus using different indices. Our results show a strong local heterogeneity in yields, underscoring the challenge of implementing index insurance in smallholder systems, and the potential benefits of low-cost yield measurement approaches that can enable more local definitions of insurance zones.
Rapid and accurate assessment of building damage in sudden-onset disasters is crucial for effective humanitarian assistance and disaster response. However, the occurrence of disasters is highly uncertain, e.g., unexpected geographic location and hazards, which challenge the conventional building damage assessment model on generalization and transferability. Unfortunately, there is little public literature on transferable building damage assessment. This is because assessing building damage using pre- and post-disaster satellite images is a complex, multi-temporal, and multi-task problem. It involves two main subtasks: building localization and damage classification, which are non-trivial to handle with generic transfer learning approaches designed for single-image and single-task problems. On the other hand, post-disaster training image availability in the target domain remains an obstacle since these generic transfer learning methods require pre-/post-disaster image pairs as target training images, resulting in a costly time window (period from obtaining post-event training image to obtaining assessment results) in disaster response. In this paper, we present a single-temporal domain adaptive semantic change detection framework, which frames domain adaptive building damage assessment and only additionally requires target pre-disaster images for adaptation training. Our framework first presents a decoupled task modeling via the equivalent form of prediction error expectations. This enables generic transfer learning methods to be used for domain adaptive building damage assessment. To fundamentally overcome the problem of post-disaster training image availability within our framework, we propose an unsupervised single-temporal change adaptation (STCA) algorithm. The main idea is “damage is everywhere”, which is motivated by the fact that building damage is a change process driven by the disaster event. We leverage target pre-disaster images and source post-disaster images to simulate such semantic change processes to provide training data, fundamentally addressing the post-disaster training image availability issue and avoiding that costly time window. The extensive experiments on global-scale and local-scale study areas suggest that our framework allows most transfer learning approaches to work well on domain adaptive building damage assessment. Our STCA achieves superior performance compared to other transfer learning approaches. More importantly, unlike other approaches that rely on target pre/post-disaster images for adaptation, it requires no target post-disaster training images. This nature significantly improves the availability of STCA in real-world disaster response for the building damage assessment model.
Small farms contribute to a large share of the productive land in developing countries. In regions such as sub-Saharan Africa, where 80% of farms are small (under 2 ha in size), the task of mapping smallholder cropland is an important part of tracking sustainability measures such as crop productivity. However, the visually diverse and nuanced appearance of small farms has limited the effectiveness of traditional approaches to cropland mapping. Here we introduce a new approach based on the detection of harvest piles characteristic of many smallholder systems throughout the world. We present HarvestNet, a dataset for mapping the presence of farms in the Ethiopian regions of Tigray and Amhara during 2020-2023, collected using expert knowledge and satellite images, totaling 7k hand-labeled images and 2k ground-collected labels. We also benchmark a set of baselines, including SOTA models in remote sensing, with our best models having around 80% classification performance on hand labelled data and 90% and 98% accuracy on ground truth data for Tigray and Amhara, respectively. We also perform a visual comparison with a widely used pre-existing coverage map and show that our model detects an extra 56,621 hectares of cropland in Tigray. We conclude that remote sensing of harvest piles can contribute to more timely and accurate cropland assessments in food insecure regions. The dataset can be accessed through https://figshare.com/s/45a7b45556b90a9a11d2, while the code for the dataset and benchmarks is publicly available at https://github.com/jonxuxu/harvest-piles
Research on yield predictions is dominated by two approaches: machine learning and process-based models. Machine learning has shown impressive results in capturing complex relationships but is often limited by data availability in agriculture. Conversely, process-based models, with over 60 years of research history, simulate crop growth processes using biophysical equations. Here, we present a method to transfer domain knowledge from the Decision Support System for Agrotechnology Transfer framework (DSSAT) using the Nwheat crop simulation process-model into neural networks and random forest for predicting wheat yield at field scale. Expanding the feature and distribution space involved simulating crop parameters and synthetic samples through the utilization of observed and historical weather recordings, as well as future climate projections. We demonstrated that neural networks can learn both general crop growth and yield processes and then effectively adapt to regional, field-specific growth patterns using synthetic and high-resolution field data. This approach boosts overall performance and reduces model error by 8 % compared to a purely data-centric model without process-knowledge transfer and solely trained on observed field data and features. Synthetic samples generated from warmer conditions were the greatest driver for improvements and we showed that the climate scenario for data generation is more important than the actual synthetic data set size. The proposed method shows the potential of combining process-based and machine-learning models, highlighting the potential to leverage the strengths of both methods in a collaborative manner.
The application of machine learning (ML) in a range of geospatial tasks is increasingly common but often relies on globally available covariates such as satellite imagery that can either be expensive or lack predictive power. Here we explore the question of whether the vast amounts of knowledge found in Internet language corpora, now compressed within large language models (LLMs), can be leveraged for geospatial prediction tasks. We first demonstrate that LLMs embed remarkable spatial information about locations, but naively querying LLMs using geographic coordinates alone is ineffective in predicting key indicators like population density. We then present GeoLLM, a novel method that can effectively extract geospatial knowledge from LLMs with auxiliary map data from OpenStreetMap. We demonstrate the utility of our approach across multiple tasks of central interest to the international community, including the measurement of population density and economic livelihoods. Across these tasks, our method demonstrates a 70% improvement in performance (measured using Pearson's $r^2$) relative to baselines that use nearest neighbors or use information directly from the prompt, and performance equal to or exceeding satellite-based benchmarks in the literature. With GeoLLM, we observe that GPT-3.5 outperforms Llama 2 and RoBERTa by 19% and 51% respectively, suggesting that the performance of our method scales well with the size of the model and its pretraining dataset. Our experiments reveal that LLMs are remarkably sample-efficient, rich in geospatial information, and robust across the globe. Crucially, GeoLLM shows promise in mitigating the limitations of existing geospatial covariates and complementing them well. Code is available on the project website: https://rohinmanvi.github.io/GeoLLM
Crop rotation has been widely used to enhance crop yields and mitigate adverse climate impacts. The existing research predominantly focuses on the impacts of crop rotation under growing season (GS) climates, neglecting the influences of non-GS (NGS) climates on agroecosystems. This oversight limits our understanding of the comprehensive climatic impacts on crop rotation and, consequently, our ability to devise effective adaptation strategies in response to climate warming. In this study, we examine the impacts of both GS and NGS climate conditions on the yield effect of the preceding crop in corn-soybean rotation systems from 1999 to 2018 in the US Midwest. Using causal forest analysis, we estimate that crop rotation increases corn and soybean yields by 0.96 and 0.22 t/ha on average, respectively. We then employ statistical models to indicate that increasing temperatures and rainfall in the NGS reduce corn rotation benefits, while warming GS enhances rotation benefits for soybeans. By 2051-2070, we project that warming climates will reduce corn rotation benefits by 6.74% under Shared Socioeconomic Pathway (SSP) 1-2.6 and 17.18% under SSP 5-8.5. For soybeans, warming climates are expected to increase rotation benefits by 8.36% under SSP 1-2.6 and 13.83% under SSP 5-8.5. Despite these diverse climate impacts on both crops, increasing crop rotation could still improve county-average yields, as neither corn nor soybean was fully rotated. If we project that all continuous corn and continuous soybeans are rotated by 2051-2070, county-average corn yields will increase by 0.265 t/ha under SSP 1-2.6 and 0.164 t/ha under SSP 5-8.5, while county-average soybean yields will gain 0.064 t/ha under SSP 1-2.6 and 0.076 t/ha under SSP 5-8.5. These findings highlight the effectiveness of crop rotation in the face of warming NGS and GS in the future and can help evaluate opportunities for adaptation.