This article presents Prithvi-EO-2.0, a new geospatial foundation model (GFM) that offers significant improvements over its predecessor, Prithvi-EO-1.0. Trained on 4.2 million global time-series samples from NASA's Harmonized Landsat and Sentinel-2 data archive at 30-m resolution, the new model incorporates temporal and location embeddings for enhanced performance across various geospatial tasks. Through extensive benchmarking with GEO-Bench, the model outperforms the previous Prithvi-EO model by 8% across a range of tasks. It also outperforms six other GFMs when benchmarked on remote sensing tasks from different domains and resolutions (i.e., from 0.1 to 15 m). The results demonstrate the versatility of the model in both classical Earth observation (EO) and high-resolution applications. Early involvement of end-users and subject matter experts (SMEs) allowed constant feedback on model and dataset design, enabling customization across diverse SME-led applications in disaster response, land cover and crop mapping, and ecosystem dynamics monitoring. Prithvi-EO-2.0 is available as an open-source model on Hugging Face and IBM TerraTorch, with additional resources on GitHub. The project exemplifies the Trusted Open Science approach embraced by all involved organizations.
Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challenging. Existing code search benchmarks focus on general software engineering tasks and fail to capture the domain-specific vocabulary and needs of scientific computing. We present a curated corpus of 5,264 high-quality, domain-classified scientific repositories spanning five NASA Science Mission Directorate divisions – Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological Physical Sciences – enriched with cleaned READMEs, extracted topics, and additional context from crawled links. Building on this corpus, we introduce two novel information retrieval benchmarks: (1) a repository search benchmark with 219 expert-curated queries designed by domain scientists, and (2) a large-scale code snippet retrieval benchmark containing 117,950 code snippets and 119,720 queries across seven programming languages. Baseline evaluations on repository search reveal significant performance variation across scientific domains. Code snippet retrieval proves equally challenging, with substantial variation driven by differing documentation practices, coding standards, and programming language conventions across scientific communities. All datasets and benchmarks are publicly released on HuggingFace to support research on scientific tool discovery.
High-quality openly-accessible machine learning (ML)-ready datasets play a foundational role in developing new artificial intelligence (AI) models or fine-tuning existing models for scientific applications such as weather and climate analysis. However, despite the growing development of new deep learning models for weather and climate, there is a scarcity of curated, pre-processed ML-ready datasets. Curating such high-quality datasets for developing new models is challenging particularly because the modality of the input data varies significantly for different downstream tasks addressing different atmospheric scales (spatial and temporal). Here we introduce WxC-Bench (Weather and Climate Bench), a multi-modal dataset designed to support the development of generalizable AI models for various downstream use-cases in weather and climate research. WxC-Bench supports examining several atmospheric processes from meso-β (20 - 200 km) scale to synoptic scales (2500 km), such as aviation turbulence, hurricane intensity and track monitoring, weather analog search, gravity wave parameterization, and natural language report generation. We provide a comprehensive description of the dataset and also present a technical validation for baseline analysis. The dataset and code to prepare the ML-ready data have been made publicly available on Hugging Face, and can be accessed using WxC-Bench Python package.
The martian atmosphere hosts dynamical phenomena ranging from planet-encircling dust storms to mesoscale orographic clouds and nocturnal low-level jets. General circulation model show capability to simulate these phenomena, but is computationally expensive at resolution needed to resolve mesoscale features. While assimilation of satellite remote sensing observation enable forecasting capabilities using such models, observation record is often sparse, short and fragmented across instrument generators. These constraints motivate the development of a data-driven foundation model for the Martian atmosphere. Foundation models live in a complex design landscape. There is an interplay between the available data, the physics of the underlying processes and corresponding developments in AI. Even though the idea of a foundation model is to address multiple use cases in a data- and compute-efficient manner, it is important to have a clear picture what applications can sensibly addressed by a single model. The purpose of this paper is to elucidate this design landscape. We discuss available data ranging from atmospheric retrievals to reanalysis datasets as well as existing physical models. Moreover, we identify a wide range of candidate downstream applications. Finally, we consider relevant recent developments in artificial intelligence (AI) that can be leveraged in this context. Here, we put a particular emphasis on AI models for atmospheric physics, data-driven approaches to data assimilation as well as methods to work in a limited data setting.
Pretraining a foundation model using MODIS observations of the earth’s atmosphere The earth and atmospheric sciences research community has an unprecedented opportunity to exploit the vast amount of data available from earth observation (EO) satellites and earth system models (ESM). Smaller and cheaper satellites with reduced operational costs have made a variety of EO data affordable, and technological advances have made the data accessible to a wide range of stakeholders, especially the scientific community (EY, 2023). The NASA ESDS program alone is expected to host 320 PB of data by 2030 (NASA ESDS, 2023). The ascent and application of artificial intelligence foundation models (FM) can be attributed to the availability of large volumes of curated data, accessibility to extensive compute resources and the maturity of deep learning architectures, especially the transformer (Bommasani et al., 2021). Developing a foundation model involves pretraining a suitable deep learning architecture with large amounts of data, often via self supervised learning (SSL) methods. The pretrained models can then be adapted to downstream tasks via fine tuning, requiring less amount of data than task-specific models. Large language models (LLM) are likely the most common type of foundation encountered by the general public. Vision transformers (ViT) are based on the LLM architecture and adapted for image and image-like data (Dosovitskiy, et. al., 2020), such as EO data and ESM simulation output. We are in the process of pretraining a ViT model for the earth’s atmosphere using a select few bands of 1-km Level-1B MODIS radiances and brightness temperatures, MOD021KM and MYD021KM from the NASA Terra and Aqua satellites respectively. We are using 200 million image chips of size 128x128 pixels. We are pretraining two ViT models of sizes 100 million and 400 million parameters respectively. The pretrained models will be finetuned for cloud classification and evaluated against AICCA. We will discuss our experiences involving data and computing, and present preliminary results. ReferencesBommasani R, Hudson DA, Adeli E, Altman R, Arora S, et al: On the opportunities and risks of foundation models. CoRR abs/2108.07258. https://arxiv.org/abs/2108.07258, 2021. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S. and Uszkoreit, J.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.Ernst & Young (EY): How can the vantage of space give you strategic advantage on Earth? https://www.ey.com/en_gl/technology/how-can-the-vantage-of-space-give-you-strategic-advantage-on-earth, 2023. Accessed 10 January 2024.Kurihana, Takuya, Elisabeth J. Moyer, and Ian T. Foster: AICCA: AI-Driven Cloud Classification Atlas. Remote Sensing 14, no. 22: 5690. https://doi.org/10.3390/rs14225690, 2022.NASA MODIS: MODIS - Level 1B Calibrated Radiances. DOI: 10.5067/MODIS/MOD021KM.061 and DOI: 10.5067/MODIS/MYD021KM.061NASA ESDS: Earthdata Cloud Evolution https://www.earthdata.nasa.gov/eosdis/cloud-evolution. Accessed 10 January 2024.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I: Attention is all you need. Adv Neural Inf Process Syst 30, 2017.
Hurricane forecasting has traditionally relied on numerical weather prediction (NWP) models. However, advancements in artificial intelligence (AI) offer new opportunities to improve forecasting accuracy. This study presents a novel evaluation of the FourCastNet model, trained on MERRA-2 and ERA5 data sets. We perform a comprehensive comparison between the FourCastNet model forecasts and those simulated by the Weather Research and Forecating (WRF) model, a NWP model, assessing both the accuracy and radial distribution of hurricane structure. This comparison provides their representation of hurricane dynamics, including differences in track prediction and intensity forecasts. Additionally, the study addresses the challenge of bias in hurricane intensity forecasts. To overcome this, this study presents a comprehensive assessment of three hurricane intensity estimation models, HxUnet, HxCNN, and HxGNN. Our results demonstrate that HxUnet consistently outperforms the other models, achieving up to a 79% reduction in maximum sustained wind speed errors and a 59% reduction in Mean Sea Level Pressure errors. This significant improvement underscores the potential of AI models to enhance the precision of hurricane intensity forecasts. This research advances the application of AI in meteorology and establishes a foundation for future studies aimed at improving hurricane prediction and mitigation efforts.
Accurate hurricane intensity estimation is critical for disaster preparedness, yet remains challenging for weather models trained on coarse-resolution datasets. This study proposes a hybrid approach that integrates NASA-IBM's Prithvi WxC model with a deep learning-based Hurricane Intensity Estimation (HIE) model. While the Prithvi WxC model excels in global atmospheric predictions, its coarse-grained outputs can struggle with precise hurricane intensity estimation. To address this, the HIE model is triggered when it identifies a hurricane in the Prithvi model output, providing corrected intensity predictions based on high-resolution data.A dataset was created for training and evaluation, consisting of 6,000 unique initial conditions from 1980 to 2024 that resulted in hurricanes across all major basins. Ground truth hurricane tracks and intensity data were obtained from the HURDAT database The training phase focused on hurricane cases from 1980 to 2000, building a foundational understanding of global hurricane characteristics. Subsequently, the model was fine-tuned with 2000–2020 data to account for basin-specific variations and improve regional accuracy. The remaining cases (2020–2024) are reserved for validation and assessment. The HIE model employs advanced deep learning techniques to refine key intensity metrics, such as maximum sustained wind speeds and central pressure. By addressing the limitations of Prithvi WxC's coarse-resolution training data, the HIE model achieves greater precision, leveraging fine-grained atmospheric and oceanographic features. This two-step framework, hurricane detection by Prithvi WxC followed by intensity refinement by the HIE model, capitalizes on the strengths of both models to deliver improved predictions.This highlights the potential of combining foundation models like Prithvi WxC with specialized deep-learning frameworks to overcome existing limitations in hurricane intensity estimation. By incorporating diverse data sources and leveraging modern machine-learning techniques, this hybrid approach bridges the gap between coarse-grained global models and the need for precise regional forecasting.
AI-based weather emulators have begun to rival the accuracy of traditional numerical solvers, for a fraction of the computational cost. The question of whether they can be reliably deployed in all use cases (e.g., for the forecast of extreme scenarios), however, is still open. We outline an ensembling strategy based on architectural variations of the Prithvi WxC foundation model (FM), highlighting the impact of each of these variations on physical accuracy and ability to capture the distributional extremes. A simple of ensemble of 100 models is sufficient to observe the complex mapping between configuration parameters and the forecast sensitivity of different atmospheric variables. We characterize some features of this mapping and connect them to the task of predicting various weather extremes.
Triggered by the realization that AI emulators can rival the performance of traditional numerical weather prediction models running on HPC systems, there is now an increasing number of large AI models that address use cases such as forecasting, downscaling, or nowcasting. While the parallel developments in the AI literature focus on foundation models – models that can be effectively tuned to address multiple, different use cases – the developments on the weather and climate side largely focus on single-use cases with particular emphasis on mid-range forecasting. We close this gap by introducing Prithvi WxC, a 2.3 billion parameter foundation model developed using 160 variables from the Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA-2). Prithvi WxC employs an encoder-decoder-based architecture, incorporating concepts from various recent transformer models to effectively capture both regional and global dependencies in the input data. The model has been designed to accommodate large token counts to model weather phenomena in different topologies at fine resolutions. Furthermore, it is trained with a mixed objective that combines the paradigms of masked reconstruction with forecasting. We test the model on a set of challenging downstream tasks namely: Autoregressive rollout forecasting, Downscaling, Gravity wave flux parameterization, and Extreme events estimation. The pretrained model with 2.3 billion parameters, along with the associated fine-tuning workflows, has been publicly released as an open-source contribution via Hugging Face.
This paper presents Prithvi-EO-2.0, a new geospatial foundation model that offers significant improvements over its predecessor, Prithvi-EO-1.0. Trained on 4.2 million global time series samples from NASA's Harmonized Landsat and Sentinel-2 data archive at 30-m resolution, the new model incorporates temporal and location embeddings for enhanced performance across various geospatial tasks. Through extensive benchmarking with GEO-Bench, the model outperforms the previous Prithvi-EO model by 8
Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled datasets through self-supervision, and then fine-tuned for various downstream tasks with small labeled datasets. This paper introduces a first-of-a-kind framework for the efficient pre-training and fine-tuning of foundational models on extensive geospatial data. We have utilized this framework to create Prithvi, a transformer-based geospatial foundational model pre-trained on more than 1TB of multispectral satellite imagery from the Harmonized Landsat-Sentinel 2 (HLS) dataset. Our study demonstrates the efficacy of our framework in successfully fine-tuning Prithvi to a range of Earth observation tasks that have not been tackled by previous work on foundation models involving multi-temporal cloud gap imputation, flood mapping, wildfire scar segmentation, and multi-temporal crop segmentation. Our experiments show that the pre-trained model accelerates the fine-tuning process compared to leveraging randomly initialized weights. In addition, pre-trained Prithvi compares well against the state-of-the-art, e.g., outperforming a conditional GAN model in multi-temporal cloud imputation by up to 5pp (or 5.7%) in the structural similarity index. Finally, due to the limited availability of labeled data in the field of Earth observation, we gradually reduce the quantity of available labeled data for refining the model to evaluate data efficiency and demonstrate that data can be decreased significantly without affecting the model's accuracy. The pre-trained 100 million parameter model and corresponding fine-tuning workflows have been released publicly as open source contributions to the global Earth sciences community through Hugging Face.
Abstract Taking the examples of Hurricane Florence (2018) over the Carolinas and Hurricane Harvey (2017) over the Texas Gulf Coast, the study attempts to understand the performance of slab, single‐layer Urban Canopy Model (UCM), and Building Environment Parameterization (BEP) in simulating hurricane rainfall using the Weather Research and Forecasting (WRF) model. The WRF model simulations showed that for an intense, large‐scale event such as a hurricane, the model quantitative precipitation forecast over the urban domain was sensitive to the model urban physics. The spatial and temporal verification using the modified Kling‐Gupta efficiency and Method for Object based Diagnostic and Evaluation in Time Domain suggests that UCM performance is superior to the BEP scheme. Additionally, using the BEP urban physics scheme over UCM for landfalling hurricane rainfall simulations has helped simulate heavy rainfall hotspots.
<p>Transformers in general have shown great promise in the sequence modeling. Recently proposed vision transformer (ViT) by Dosovitskiy et al. has shown optimal performance in image recogining [1]. Fourier Neural operator based token mixer transformers keeping ViT as backbone was proposed by Guibas et. al. has been used for predicting wind and precipitation on ERA5 dataset[2,3]. Following the previous work, we trained the Fourcastnet from scratch on the MERRA2 data set with 3 verticle levels (z450, z500, z550) and 11 variables (adding u, v, and temp). We trained on data from 2005 to 2015 and made predictions by providing the initial conditions from 2017. The prediction was made for 7 days in advance. For the first 24 hours model prediction, mean correlation was 0.998. Root mean squared error (RMSE) 6 hours prediction was 8.779 and for 24 hours was 19.581 on a scale range of -575.6 to 330.6. The model was further tested on 11 variables on the same training data to evaluate prediction of major events like Hurricane. Initial condition for category 5 Hurricane Sep 28, 2016 &#8211; Oct 10, 2016 was given to the model. The model was able to predict the hurricane for 18 hours. Further work will be done in order to tune to model and increase more environment variables from MERRA2 to make the prediction more robust and for a longer period.</p><p>References:<br>1. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold,<br>G., Gelly, S. and Uszkoreit, J., 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv<br>preprint arXiv:2010.11929.<br>2. Guibas, J., Mardani, M., Li, Z., Tao, A., Anandkumar, A. and Catanzaro, B., 2021, September. Efficient Token Mixing for<br>Transformers via Adaptive Fourier Neural Operators. In International Conference on Learning Representations<br>3. Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A., Mardani, M., Kurth, T., Hall, D., Li, Z.,<br>Azizzadenesheli, K. and Hassanzadeh, P., 2022. Fourcastnet: A global data-driven high-resolution weather model using<br>adaptive fourier neural operators. arXiv preprint arXiv:2202.11214.</p>
Ocean warming influences tropical cyclone (TC) destructive parameters and hence their guidance certainly helps the disaster management agencies to reduce the damage. The present study modelled the sensitivity of TC size, intensity, rainfall, and destructive potential parameters under various ocean warming conditions. A recent extremely severe cyclonic storm FANI (2019) over the Bay of Bengal is chosen for this purpose. The 9-km grid-spacing Weather Research and Forecasting (WRF) model simulations are performed by altering the default sea surface temperature (SST) by-1 degrees C, +1 degrees C, +2 degrees C and, +3 degrees C respectively along with the control run. The model simulations revealed that ocean warming causes the FANI cyclone to turn northeastward direction. The observed changes of tangential wind speed due to large sea surface enthalpy fluxes associated with ocean warming result in TC size changes and then guide the cyclone to northeastward movement. The radius of 34-knot wind (R34) is more sensitive to the SST warming compared to the radius of maximum winds. The modelled R34 values ranged between similar to 180-600 km against observations (similar to 100-400 km) during the TC life period. The increased tangential wind speeds (similar to 50-60 m s(-1)) and convective updrafts (similar to 1.3-1.5 m s(-1)) in the TC area are responsible for TC intensity and size changes. A linear and exponential growth is seen for the TC destructive potential indicators those estimated using TC intensity and R34 values. The SST increase could result in peak heavy rainfall (>65 mm day(-1)) in the TC inner-core region, especially the rear sector. The categorical rainfall distribution analysis also proved that heavy rainfall areas extend to greater distances (>300 km) around the TC center with SST warming. The study helps in assessing the hydro-meteorological destruction associated with the TCs in the future warming climate.
The present study examines the climatological characteristics and possible triggers of rapid intensification (RI) of tropical cyclones (TCs) during 1990–2019 over the North Indian Ocean (NIO). RI is defined as an increase in maximum sustained surface wind speed of 30 knots (15.4 m⋅s−1) or more in a 24 hr duration. In the NIO basin, the threshold of 24 hr intensity change represents the 93rd percentile using a 3‐min sustained wind, while, it is 95th percentile in the Atlantic basin, where 1‐min sustained wind is used. A total of 46 TCs (~38% of the total TCs) have exhibited RI at an average rate of 1–2 TCs year−1 over the NIO. A significant increase of RI‐TCs is seen from the year 2000 onwards over the region. The maximum RI‐TCs occurrence is found in the post‐monsoon season. The majority of TCs (~48%) undergo the RI phase within a 12–24 hr time during the depression stage. About ~35% (26%) of the TCs are retained in the RI phase for the duration of at least 24 hr (36 hr). Most of the RI‐TCs move northwestward (38%) and westward (31%) direction 6 hr before the RI onset with slow/normal translation speeds. During the RI phase, ~72% of TCs travels a distance of ~150–450 km. The TC inner‐core region receives heavy rainfall, and a ~ 3 mm⋅hr−1 increment is noticed 12 hr before the RI to the RI onset. Most of the RI‐TCs made landfall over eastern states of India showing the vulnerability of the regions. Composite analysis demonstrated that higher precipitable water (~55 mm), surface flux (500 W⋅m−2), cyclone heat potential (50–60 kJ⋅cm−2) in the moderate shear (6–8 m⋅s−1) environment favours the RI process. This study highlights the significance of RI‐TCs and triggering conditions for the RI onset over the NIO region.
This study investigates the influence of LULC representation on surface meteorological conditions associated with monsoon depressions (MDs) in the Advanced Research Weather Research and Forecasting (WRF) model. A total of 18 MDs were considered during 2007–2018, with a life period of at least two days. Two simulations are performed at a 5-km grid-spacing, ingesting the LULC from the United States Geological Survey (USGS) and National Remote Sensing Centre (NRSC) for each MD case. The urban area has been increased from 0.13% in USGS to 1.21% in NRSC over the north-central and east coastal states, reducing the soil moisture (SM) errors by 0.015 m3 m−3 (25%) in the NRSC run. The SM in the NRSC run has a correlation of 0.53, which is 15% higher than that in the USGS run. The NRSC has a larger forest area (22%) than the USGS (7%) over the north-western parts and some parts of Maharashtra, which helps increase latent heat flux (LHF) by 30 Wm−2. Verifying with ERA analysis, the USGS and the NRSC simulations underestimate the LHF in most parts of India and overestimate in orographic areas. The NRSC-simulation shows fewer improvements (<5%) in LHF across the central and southern areas, unlike in the East and West parts (>13%). The surface temperature, moisture, and rainfall have been noticeably improved in the NRSC run for day-2 and day-3 simulation, unlike in USGS, while they are comparable in day-1. The spatial error of rainfall increases with forecast length, with NRSC having 10–15% less error than the USGS run. Further, the NRSC run could identify the rainfall peaks of 3 mm h−1 with a mean error of 1.5 mm h−1 as compared to 2.2 mm h−1 error in the USGS. This study demonstrates the positive impact of NRSC-developed LULC in the MD simulations over India.
Soil moisture and temperature (SM and ST) have been identified for modeling of extreme weather and hydrological processes. The coarser resolution global analyses are limited in capturing realistic heterogeneity. This study focuses on evaluating regional land surface conditions developed from a high-resolution (4 km grid spacing) land data assimilation system (HRLDAS) over India from 2000 to 2013 against in situ and global analyses. Global analyses such as the European Space Agency Climate Change Initiative (ESACCI), Moderate Resolution Imaging Spectroradiometer (MODIS), Climate Forecasting System (CFS), and Global Land Data Assimilation System (GLDAS) have been considered to assess the credibility of the regional analysis. The regional SM from the HRLDAS is superior to global and satellite products, particularly in the orography (altitude > 300 m) regions followed by the plane regions (altitude ≤ 300 m). The probability distribution function (PDF) indicates that the regional SM and satellite analysis exhibited less error (~ 0.02 m3 m−3 at ~ 28%) in the plane regions. The regional SM analysis in the orography regions is reliable (0.015 m3 m−3 at 28% frequency) with a high equitable threat score (~ 0.6) compared to other analyses. The HRLDAS is consistently superior for soil temperature (ST) to other global analyses. The mean diurnal variation of HRLDAS-ST is close to in situ observation. The HRLDAS performs better for spatial representation of SM and ST for different months and monsoon seasons. The improved representation of land conditions from the HRLDAS could provide a realistic distribution of latent and sensible heat fluxes when compared with other global products. This study demonstrates the value of high-resolution regional analyses and recommends usefulness in hydrological applications.
The Hurricane Weather Research and Forecasting (HWRF) is increasingly becoming known for its better performance in simulating tropical cyclones (TCs) over different basins globally, and the Advanced Research version of the Weather Research and Forecasting (WRF) model is widely used for the same in both research and operational settings. These two models are now operational at India Meteorological Department (IMD) for TC predictions over the North Indian Ocean. The near real-time forecast of HWRF and WRF models is evaluated in a quasi-operational setup based on 62 forecast cases from 10 recent TCs during 2013–2017 over the Bay of Bengal. Multi-satellite estimated winds are used to compare initial and simulated vortex structures, and HWRF has been found to be better. The track prediction is comparable in both the models for shorter forecast lengths (up to 30 h), while the HWRF is skillful (by 27%) for longer forecast, thus leading to a better estimation of landfall position and time. HWRF is significantly better (error < 10 knots) in intensity prediction compared to the WRF model (~ 15 knots). Unlike the WRF model, the HWRF model produced an improved vertical structure of dynamic and thermodynamic processes and could be attributed to the state-of-the-art vortex initialization and relocation method, multiscale interaction and high resolution. The model-predicted vortex structures also supported the credibility of HWRF system. This work highlights the need for having improved initialization and high resolution for TC predictions over the region.
Several recent papers have investigated different challenges in applying machine learning (ML) techniques to Earth science problems. The challenges listed range from interpretability of the results to computational demand to data issues. In this paper, we focus on specific challenges listed in the review papers that are centered around training data, as the size of training data is important in applying deep learning (DL) techniques. We are in the process of conducting a literature survey to better understand these challenges as well as to understand any trends. As part of this survey, our review has encompassed Earth science papers from AGU, AMS, IEEE and SPIE journals covering the last ten years and focused on papers that utilize supervised ML techniques. Our initial survey results show some interesting findings. The use of supervised machine learning techniques in Earth science research has increased significantly in the last decade. The number of atmospheric science papers (i.e., from AMS journals) using ML approaches has increased by over 40%. Across all of Earth science even larger changes have occurred, including a >90% increase in AGU papers and a >10-fold increase in IEEE papers using ML. We also conducted a deep dive into all the papers from AGU journals and uncovered interesting findings. There is a prevalence of the use of supervised ML in certain sub-disciplines within Earth science. The biogeoscience and land surface research communities lead in this area: over 20% of papers published in Global Biogeochemical Cycles, JGR Biogeosciences, JGR Earth Surface, and Water Resources Research use supervised ML techniques, including over 35% of the papers in JGR Biogeosciences. The availability of labeled training data in Earth science is reflected in the number of training samples used in supervised analysis. In the papers we surveyed, most ML algorithms were trained using small (i.e. hundreds of labeled) samples. However, for some applications using model output or large, established datasets, the number of training data ranged several orders of magnitude greater. In this presentation, we will describe our findings from the literature survey. We will also list recommendations for the science community to address the existing challenges around training data.
Recent review papers have discussed the opportunities and challenges of applying machine learning (ML) techniques to Earth science data. A common challenge cited in these papers is the lack of labeled training data. A literature review of Earth science papers over the last 10 years demonstrates that while there is rapid adoption of ML, particularly in biogeoscience and land surface research, the training datasets typically contain only hundreds of samples. This lack of training data limits the use of deep learning algorithms, which require larger volumes of labeled data. In situ training data are most frequently used in almost all domains, followed by model output and satellite data. The atmosphere and solid Earth domains use the largest training datasets, an order of magnitude larger than in biogeoscience papers. Random forest is the most commonly applied ML algorithm in all domains except atmospheric science and biogeoscience, which more frequently use fully connected neural networks.