Tropical cyclones are among the most dangerous and costly weather phenomena, but forecasting them remains challenging. Here we introduce WeatherNext Cyclones (WN-C), an artificial intelligence (AI)-based operational weather model that produces state-of-the-art ensemble forecasts of the track, intensity and size of tropical cyclones worldwide. Trained on a combination of global analysis data1 and a global database of historical tropical cyclones2,3, WN-C generates large ensembles of possible global weather and cyclone scenarios extending 15 days into the future. When evaluated on tropical cyclones from 2023 to 2025, the track, intensity and wind-radius predictions from WN-C offer an average lead-time advantage of 1 day or more over leading operational models—an improvement in accuracy comparable to the progress seen in the last decade of operational development. We achieved these results using inputs that are orders of magnitude coarser than regional models, suggesting that high resolution is not a strict prerequisite for state-of-the-art intensity forecasting and that these coarse atmospheric data contain more intensity signal than has been previously recognized. Including predictions from WN-C in a weighted-average consensus ensemble improves its skill substantially. The scalability of WN-C enables ensembles of up to 1,000 members, which are better at capturing rare events than conventional 50-member ensembles. By providing advanced operational ensemble guidance to human forecasters, this work represents a step change towards more reliable and timely forecasts and warnings that can help to protect lives and to mitigate the devastating effects of tropical cyclones. An AI forecasting system called WeatherNext Cyclones can predict the track, intensity (maximum wind speed) and size of tropical cyclones with an accuracy gain over leading operational models comparable to a decade of progress in numerical forecasting.
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.
Conventional studies of subseasonal-to-seasonal sea ice variability across scales have relied upon computationally expensive physics-based models solving systems of differential equations. IceNet, a deep learning-based sea ice forecasting model under development since 2021, has proven competitive to such state-of-the-art physics-based models, capable of generating daily 25 km resolution forecasts of sea ice concentration across the Arctic and Antarctic at a fraction of the computational cost once trained. Yet, these IceNet forecasts leave room for improvement through three main weaknesses. First, the forecasts exhibit physically unrealistic spatial and temporal blurring characteristic of deep learning methods trained under mean loss objectives. Second, the use of 25 km scale OSISAF data renders local forecasts along coastal regions and in regions surrounding maritime vessels inconclusive. Third, the sole provision of sea ice concentration in forecasts leaves questions about other critical ice properties such as thickness unanswered. We present preliminary results addressing these three challenges, turning to deep generative models to capture forecast uncertainty and improve spatial sharpness; leveraging 3 and 6 km scale AMSR-2 sea ice products to improve spatial resolution; and incorporating auxiliary datasets, chiefly thickness, into the training and inference pipeline to produce multivariate forecasts of sea ice properties beyond simple sea ice concentration. We seek feedback for improvement and hope continued development of IceNet can help answer key scientific questions surrounding the state of sea ice in our changing polar climates.
Weather prediction is critical for a range of human activities, including transportation, agriculture and industry, as well as for the safety of the general public. Machine learning transforms numerical weather prediction (NWP) by replacing the numerical solver with neural networks, improving the speed and accuracy of the forecasting component of the prediction pipeline1-6. However, current models rely on numerical systems at initialization and to produce local forecasts, thereby limiting their achievable gains. Here we show that a single machine learning model can replace the entire NWP pipeline. Aardvark Weather, an end-to-end data-driven weather prediction system, ingests observations and produces global gridded forecasts and local station forecasts. The global forecasts outperform an operational NWP baseline for several variables and lead times. The local station forecasts are skilful for up to ten days of lead time, competing with a post-processed global NWP baseline and a state-of-the-art end-to-end forecasting system with input from human forecasters. End-to-end tuning further improves the accuracy of local forecasts. Our results show that skilful forecasting is possible without relying on NWP at deployment time, which will enable the realization of the full speed and accuracy benefits of data-driven models. We believe that Aardvark Weather will be the starting point for a new generation of end-to-end models that will reduce computational costs by orders of magnitude and enable the rapid, affordable creation of customized models for a range of end users.
Every autumn on the south coast of Victoria Island (Nunavut, Canada), endangered Dolphin and Union (DU) caribou ( Rangifer tarandus groenlandicus x pearyi ) wait for sea ice to form before continuing their southwards migration to the mainland. Delayed freeze‐up, less stable ice conditions and ice‐breaking by vessels are putting migrating caribou at risk, but unpredictable freeze‐up times pose challenges for conservation planning. Having early warning of when the caribou sea ice crossing is likely to take place could guide more targeted measures (e.g., ice‐breaking vessel management). In this case study, we use a multi‐stakeholder approach to explore the potential of using observed and forecast sea ice concentration (SIC) to predict when DU caribou are likely to cross the sea ice. We examine links between caribou movement records and coincident satellite observations of SIC collected between 1996–2005 and 2015–2019. We establish probabilistic “percent‐crossed” metrics to convert SIC freeze‐up profiles into anticipated sea ice crossing‐start date ranges and maps. Finally, we assess the potential of using IceNet, an AI‐based 25 km resolution SIC forecast model, to predict these crossing‐start ranges in 2020–2022. We identify a clear link between SIC freeze‐up profiles and crossing‐start times, with median SIC reaching 98.8% (IQR = 94.1%, 100%) when caribou start their crossings. Our percent‐crossed metrics are effective in converting SIC records into crossing‐start date maps which can guide human experts. IceNet results show promise, predicting crossing‐start ranges comparable to those observed in 2022 up to three weeks before the first observed sea ice crossing. In 2021, IceNet's predicted ranges are systematically early, but improve between three‐ to one‐week lead times. Practical implication : AI sea ice forecasts could provide early warning of DU caribou sea ice crossing times, informing mitigation of ice‐breaking vessels and providing a blueprint applicable to other ice‐dependent species. Our case study contributes practical considerations, limitations and areas for future research to drive innovation in this emerging field forward. Ultimately, forecasts could be integrated into human‐expert centred decision‐support tools, guiding dynamic conservation and management for Arctic species.
Machine learning (ML)-based weather models have rapidly risen to prominence due to their greater accuracy and speed than traditional forecasts based on numerical weather prediction (NWP), recently outperforming traditional ensembles in global probabilistic weather forecasting. This paper presents FGN, a simple, scalable and flexible modeling approach which significantly outperforms the current state-of-the-art models. FGN generates ensembles via learned model-perturbations with an ensemble of appropriately constrained models. It is trained directly to minimize the continuous rank probability score (CRPS) of per-location forecasts. It produces state-of-the-art ensemble forecasts as measured by a range of deterministic and probabilistic metrics, makes skillful ensemble tropical cyclone track predictions, and captures joint spatial structure despite being trained only on marginals.
Weather forecasts are fundamentally uncertain, so predicting the range of probable weather scenarios is crucial for important decisions, from warning the public about hazardous weather to planning renewable energy use. Traditionally, weather forecasts have been based on numerical weather prediction (NWP)1, which relies on physics-based simulations of the atmosphere. Recent advances in machine learning (ML)-based weather prediction (MLWP) have produced ML-based models with less forecast error than single NWP simulations2,3. However, these advances have focused primarily on single, deterministic forecasts that fail to represent uncertainty and estimate risk. Overall, MLWP has remained less accurate and reliable than state-of-the-art NWP ensemble forecasts. Here we introduce GenCast, a probabilistic weather model with greater skill and speed than the top operational medium-range weather forecast in the world, ENS, the ensemble forecast of the European Centre for Medium-Range Weather Forecasts4. GenCast is an ML weather prediction method, trained on decades of reanalysis data. GenCast generates an ensemble of stochastic 15-day global forecasts, at 12-h steps and 0.25 degrees latitude-longitude resolution, for more than 80 surface and atmospheric variables, in 8 min. It has greater skill than ENS on 97.2% of 1,320 targets we evaluated and better predicts extreme weather, tropical cyclone tracks and wind power production. This work helps open the next chapter in operational weather forecasting, in which crucial weather-dependent decisions are made more accurately and efficiently.
Weather forecasting is critical for a range of human activities including transportation, agriculture, industry, as well as the safety of the general public. Machine learning models have the potential to transform the complex weather prediction pipeline, but current approaches still rely on numerical weather prediction (NWP) systems, limiting forecast speed and accuracy. Here we demonstrate that a machine learning model can replace the entire operational NWP pipeline. Aardvark Weather, an end-to-end data-driven weather prediction system, ingests raw observations and outputs global gridded forecasts and local station forecasts. Further, it can be optimised end-to-end to maximise performance over quantities of interest. Global forecasts outperform an operational NWP baseline for multiple variables and lead times. Local station forecasts are skillful up to ten days lead time and achieve comparable and often lower errors than a post-processed global NWP baseline and a state-of-the-art end-to-end forecasting system with input from human forecasters. These forecasts are produced with a remarkably simple neural process model using just 8 less compute than existing NWP and hybrid AI-NWP methods. We anticipate that Aardvark Weather will be the starting point for a new generation of end-to-end machine learning models for medium-range forecasting that will reduce computational costs by orders of magnitude and enable the rapid and cheap creation of bespoke models for users in a variety of fields, including for the developing world where state-of-the-art local models are not currently available.
Machine learning (ML)-based weather models have recently undergone rapid improvements. These models are typically trained on gridded reanalysis data from numerical data assimilation systems. However, reanalysis data comes with limitations, such as assumptions about physical laws and low spatiotemporal resolution. The gap between reanalysis and reality has sparked growing interest in training ML models directly on observations such as weather stations. Modelling scattered and sparse environmental observations requires scalable and flexible ML architectures, one of which is the convolutional conditional neural process (ConvCNP). ConvCNPs can learn to condition on both gridded and off-the-grid context data to make uncertainty-aware predictions at target locations. However, the sparsity of real observations presents a challenge for data-hungry deep learning models like the ConvCNP. One potential solution is 'Sim2Real': pre-training on reanalysis and fine-tuning on observational data. We analyse Sim2Real with a ConvCNP trained to interpolate surface air temperature over Germany, using varying numbers of weather stations for fine-tuning. On held-out weather stations, Sim2Real training substantially outperforms the same model architecture trained only with reanalysis data or only with station data, showing that reanalysis data can serve as a stepping stone for learning from real observations. Sim2Real could thus enable more accurate models for weather prediction and climate monitoring.
Weather forecasts are fundamentally uncertain, so predicting the range of probable weather scenarios is crucial for important decisions, from warning the public about hazardous weather, to planning renewable energy use. Here, we introduce GenCast, a probabilistic weather model with greater skill and speed than the top operational medium-range weather forecast in the world, the European Centre for Medium-Range Forecasts (ECMWF)'s ensemble forecast, ENS. Unlike traditional approaches, which are based on numerical weather prediction (NWP), GenCast is a machine learning weather prediction (MLWP) method, trained on decades of reanalysis data. GenCast generates an ensemble of stochastic 15-day global forecasts, at 12-hour steps and 0.25 degree latitude-longitude resolution, for over 80 surface and atmospheric variables, in 8 minutes. It has greater skill than ENS on 97.4 predicts extreme weather, tropical cyclones, and wind power production. This work helps open the next chapter in operational weather forecasting, where critical weather-dependent decisions are made with greater accuracy and efficiency.
Conditional neural processes (CNPs; Garnelo et al., 2018a) are attractive meta-learning models which produce well-calibrated predictions and are trainable via a simple maximum likelihood procedure. Although CNPs have many advantages, they are unable to model dependencies in their predictions. Various works propose solutions to this, but these come at the cost of either requiring approximate inference or being limited to Gaussian predictions. In this work, we instead propose to change how CNPs are deployed at test time, without any modifications to the model or training procedure. Instead of making predictions independently for every target point, we autoregressively define a joint predictive distribution using the chain rule of probability, taking inspiration from the neural autoregressive density estimator (NADE) literature. We show that this simple procedure allows factorised Gaussian CNPs to model highly dependent, non-Gaussian predictive distributions. Perhaps surprisingly, in an extensive range of tasks with synthetic and real data, we show that CNPs in autoregressive (AR) mode not only significantly outperform non-AR CNPs, but are also competitive with more sophisticated models that are significantly more computationally expensive and challenging to train. This performance is remarkable given that AR CNPs are not trained to model joint dependencies. Our work provides an example of how ideas from neural distribution estimation can benefit neural processes, and motivates research into the AR deployment of other neural process models.
Environmental sensors are crucial for monitoring weather conditions and the impacts of climate change. However, it is challenging to place sensors in a way that maximises the informativeness of their measurements, particularly in remote regions like Antarctica. Probabilistic machine learning models can suggest informative sensor placements by finding sites that maximally reduce prediction uncertainty. Gaussian process (GP) models are widely used for this purpose, but they struggle with capturing complex non-stationary behaviour and scaling to large datasets. This paper proposes using a convolutional Gaussian neural process (ConvGNP) to address these issues. A ConvGNP uses neural networks to parameterise a joint Gaussian distribution at arbitrary target locations, enabling flexibility and scalability. Using simulated surface air temperature anomaly over Antarctica as training data, the ConvGNP learns spatial and seasonal non-stationarities, outperforming a non-stationary GP baseline. In a simulated sensor placement experiment, the ConvGNP better predicts the performance boost obtained from new observations than GP baselines, leading to more informative sensor placements. We contrast our approach with physics-based sensor placement methods and propose future steps towards an operational sensor placement recommendation system. Our work could help to realise environmental digital twins that actively direct measurement sampling to improve the digital representation of reality.
Ice cores record crucial information about past climate. However, before ice core data can have scientific value, the chronology must be inferred by estimating the age as a function of depth. Under certain conditions, chemicals locked in the ice display quasi-periodic cycles that delineate annual layers. Manually counting these noisy seasonal patterns to infer the chronology can be an imperfect and time-consuming process, and does not capture uncertainty in a principled fashion. In addition, several ice cores may be collected from a region, introducing an aspect of spatial correlation between them. We present an exploration of the use of probabilistic models for automatic dating of ice cores, using probabilistic programming to showcase its use for prototyping, automatic inference and maintainability, and demonstrate common failure modes of these tools.
Deploying environmental measurement stations can be a costly and time-consuming procedure, especially in remote regions that are difficult to access, such as Antarctica. Therefore, it is crucial that sensors are placed as efficiently as possible, maximising the informativeness of their measurements. This can be tackled by fitting a probabilistic model to existing data and identifying placements that would maximally reduce the model’s uncertainty. The models most widely used for this purpose are Gaussian processes (GPs; Williams and Rasmussen, 2006). However, designing a GP covariance which captures the complex behaviour of non-stationary spatiotemporal data is a difficult task. Further, the computational cost of GPs makes them challenging to scale to large environmental datasets. In this work, we explore using a convolutional Gaussian neural process (ConvGNP; Bruinsma et al., 2021; Markou et al., 2022) to address these issues. A ConvGNP is a meta-learning model that uses neural networks to parameterise a GP predictive. Our model is data-driven, flexible, efficient, and permits multiple input predictors of gridded or scattered modalities. Using simulated surface air temperature fields over Antarctica as ground truth, we show that a ConvGNP significantly out-performs a non-stationary GP baseline in terms of predictive performance. We then use the ConvGNP in an Antarctic sensor placement toy experiment, yielding promising results.
The dataset and configuration files are for a demonstrator notebook in the Environmental AI book, https://acocac.github.io/environmental-ai-book/welcome.html. The files were generated using the IceNet source code, https://github.com/tom-andersson/icenet-paper. The description of each file as follows: - 2021_09_03_1300_icenet_demo.json: - dataset1.zip: - siconca_EASE.nc:
Arctic sea ice forecasting is a major scientific effort with fundamental challenges at play. To address such challenges, we have developed a physics-informed, data-driven sea ice forecasting system, IceNet, which outperformed a leading dynamical model (ECMWF SEAS5) in monthly-averaged forecasts of pan-Arctic sea ice concentration. IceNet adopted a U-Net deep learning architecture and was trained on over 2,000 years of CMIP6 climate simulation data. Despite its state-of-the-art seasonal forecasting skill at lead times of 2-6 months, IceNet has two main limitations. First, it could not outperform the dynamical model in short-range (1-month) forecasts. This is partly caused by IceNet operating on monthly-averages, which smears the initial conditions and weather phenomena that can dominate predictability at short time scales. Second, IceNet is afflicted by the ‘spring predictability barrier’ that affects all long range forecasts of summer. This predictability barrier arises primarily due to the importance of melt-season ice thickness conditions on summer sea ice. Here we present our early findings from IceNet2, which attempts to alleviate these issues by operating on daily-averages and including sea ice thickness as an input variable. IceNet2 paves the way for our efforts to aid the Arctic conservation community by developing the first public, operational sea ice forecasting AI.
Anthropogenic warming has led to an unprecedented year-round reduction in Arctic sea ice extent. This has far-reaching consequences for indigenous and local communities, polar ecosystems, and global climate, motivating the need for accurate seasonal sea ice forecasts. While physics-based dynamical models can successfully forecast sea ice concentration several weeks ahead, they struggle to outperform simple statistical benchmarks at longer lead times. We present a probabilistic, deep learning sea ice forecasting system, IceNet. The system has been trained on climate simulations and observational data to forecast the next 6 months of monthly-averaged sea ice concentration maps. We show that IceNet advances the range of accurate sea ice forecasts, outperforming a state-of-the-art dynamical model in seasonal forecasts of summer sea ice, particularly for extreme sea ice events. This step-change in sea ice forecasting ability brings us closer to conservation tools that mitigate risks associated with rapid sea ice loss.
In Austral summer 2016/2017, the sea ice extent (SIE) in the Weddell Sea dropped to a near-record value in the satellite era (1.88 x 10(6) km(2)), a large negative seasonal anomaly that persisted in an unprecedented fashion for the following three summers. Various atmospheric and oceanic factors played a part in the change. Ice loss started in September 2016 when the northern Weddell Sea experienced westerly winds of record strength, advecting multiyear sea ice from the region. In late 2016, a polynya over Maud Rise contributed to low SIE over the eastern Weddell Sea. With extensive areas of open water early in the summer, upper ocean temperatures increased by similar to 0.5 degrees C, with the anomalies persisting in subsequent years. The reappearance of the Maud Rise polynya in 2017, high ocean temperatures, and storms of record depth kept the summer SIE low.
Over recent decades, the Arctic has warmed faster than any region on Earth. The rapid decline in Arctic sea ice extent (SIE) is often highlighted as a key indicator of anthropogenic climate change. Changes in sea ice disrupt Arctic wildlife and indigenous communities, and influence weather patterns as far as the mid-latitudes. Furthermore, melting sea ice attenuates the albedo effect by replacing the white, reflective ice with dark, heat-absorbing melt ponds and open sea, increasing the Sun’s radiative heat input to the Arctic and amplifying global warming through a positive feedback loop. Thus, the reliable prediction of sea ice under a changing climate is of both regional and global importance. However, Arctic sea ice presents severe modelling challenges due to its complex coupled interactions with the ocean and atmosphere, leading to high levels of uncertainty in numerical sea ice forecasts. Deep learning (a subset of machine learning) is a family of algorithms that use multiple nonlinear processing layers to extract increasingly high-level features from raw input data. Recent advances in deep learning techniques have enabled widespread success in diverse areas where significant volumes of data are available, such as image recognition, genetics, and online recommendation systems. Despite this success, and the presence of large climate datasets, applications of deep learning in climate science have been scarce until recent years. For example, few studies have posed the prediction of Arctic sea ice in a deep learning framework. We investigate the potential of a fully data-driven, neural network sea ice prediction system based on satellite observations of the Arctic. In particular, we use inputs of monthly-averaged sea ice concentration (SIC) maps since 1979 from the National Snow and Ice Data Centre, as well as climatological variables (such as surface pressure and temperature) from the European Centre for Medium-Range Weather Forecasts reanalysis (ERA5) dataset. Past deep learning-based Arctic sea ice prediction systems tend to overestimate sea ice in recent years - we investigate the potential to learn the non-stationarity induced by climate change with the inclusion of multi-decade global warming indicators (such as average Arctic air temperature). We train the networks to predict SIC maps one month into the future, evaluating network prediction uncertainty by ensembling independent networks with different random weight initialisations. Our model accounts for seasonal variations in the drivers of sea ice by controlling for the month of the year being predicted. We benchmark our prediction system against persistence, linear extrapolation and autoregressive models, as well as September minimum SIE predictions from submissions to the Sea Ice Prediction Network's Sea Ice Outlook. Performance is evaluated quantitatively using the root mean square error and qualitatively by analysing maps of prediction error and uncertainty.