Smoke-related PM2.5 is the primary air quality concern during the summer in Alaska, yet accurate forecasting remains a major challenge. In this study, we use machine learning (ML) techniques to improve smoke forecasts from NOAA's HRRR-Smoke model. We find that the model underestimates surface PM2.5 by a factor of up to five during the wildfire season in Alaska. We evaluated Random Forest (RF), one-dimensional, and two-dimensional convolutional neural network (CNN1D and CNN2D) models. Among them, CNN1D performed the best, reducing the underestimation factor to two or less. Analysis of the relationships between key predictors, such as Surface and Vertical Smoke and calibrated PurpleAir observations, suggests that errors in the vertical distribution of smoke are a primary source of underestimating bias. Atmospheric sounding data further show that the HRRR-Smoke model fails to capture daytime temperature inversion layers during wildfire events. This bias is likely caused by missing fire radiative power (FRP) detections under heavy smoke or cloudy conditions, which leads to low smoke concentrations and under-represented radiation feedback necessary to maintain near-surface inversion. Although accurately representing physical and chemical processes in models remains highly challenging, our results demonstrate ML offers an effective approach to improving daily surface PM2.5 forecasts in wildfire-prone regions like Alaska.
Abstract Fire radiative power (FRP) detection of wildland fires is important in calculating a variety of variables, such as smoke emissions, which are necessary for fire weather and air quality simulations. FRP is derived from satellite observations of radiative energy emitted by fires. In this work, FRP values collected by geostationary and polar-orbiting satellites in 2018 are used in combination with weather forecast model outputs to derive a next-hour forecasted hourly FRP value unique to that model grid point. Random forest (RF) models were trained and applied to multiday burning fires with inputs including day-before average FRP satellite values, hour of day, location, temperature, wind speed, and relative humidity to produce a next-hour FRP value on the Rapid Refresh (RAP) model resolution. Models were trained on separate satellite FRP inputs given the difference in satellite sensor resolutions: one used Geostationary Operational Environmental Satellite (GOES) FRP inputs and the other used a collection of polar-orbiting satellite FRP values. Overall, mean absolute errors (MAEs) for all RF models were lower than those for the current persistence approach method used by the High-Resolution Rapid Refresh (HRRR). The main lesson learned from this paper is that a simple RF model is a good alternative to the persistence method and provides a uniquely forecasted hourly FRP value based on weather and time/location. Significance Statement This work introduces a type of machine learning model, random forest (RF), to the current method in which hourly fire radiative power (FRP) is modeled using a day-before average FRP value and a hour of day value. By using an RF, numerical weather model information is used as an input with both the day-before average FRP value and the time of day value to create a grid point and weather-specific hourly FRP forecast.
Smoke from spring boreal wildfires increasingly impacts eastern North America, exposing breeding birds to hazardous air pollution that may impact reproductive outcomes. Using data collected in 2018-2025 from 70,979 monitored nests of four widespread cavity-nesting songbirds, we found strong evidence that smoke greatly delays egg laying and can extend incubation and nestling duration. We further found that while smoke is associated with increased clutch sizes, in some species smoke exposure strongly decreases hatching or fledging success. Our results demonstrate that extreme smoke can have wide-ranging impacts on breeding birds, from altering phenology to impacting fitness. While the exact mechanisms underlying these results remain elusive, the full suite of effects suggests that modifications to adult behavior under smoky conditions is the most likely cause. As fire regimes shift, birds and other wildlife are at greater risk of exposure to toxic smoke during the breeding season, which may further exacerbate the biodiversity crisis.
The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes - including low-level jets, wind ramps, extreme wind veer, and weak wind conditions - based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.
Climate change has led to an increase in the number and size of wildfires in western North America, and their emissions of particulate matter and reactive trace gases threaten to offset otherwise improving air quality. Most air quality models do not dynamically update the chemical composition of biomass burning emissions with each model time step. However, laboratory, field, and satellite remote sensing observations indicate that the chemical composition of wildfire emissions changes with evolving combustion conditions (i.e. flaming vs. smoldering). We have previously shown that the Tropospheric Monitoring Instrument (TROPOMI) can be used to observe day-to-day changes in the chemical composition of wildfire emissions, during the transition from flaming to smoldering combustion. Here, we present the use of the Hourly Wildfire Potential index (HWP) to predict the impact of meteorological changes on wildfire activity, combined with TROPOMI observations of enhanced nitrogen dioxide (NO2) emissions compared with that of carbon monoxide (CO) ( ∆ NO2/ ∆ CO), to parameterize how biomass burning nitrogen oxide (NOx) emissions change with evolving combustion conditions. We implemented this parameterization in the High-Resolution Rapid Refresh model coupled with Chemistry (HRRR-Chem), developed by NOAA Global Systems Laboratory. HRRR-Chem is an experimental and high-resolution (3x3 km2$) atmospheric chemical transport model for the U.S. based on the Weather Research and Forecasting model coupled with Chemistry (WRF-Chem). We compared modeled smoke and trace gases with airborne measurements to evaluate our model. Overall, we found that modifying biomass burning NOx emissions based on HWP significantly improved predicted ozone concentrations in smoke plumes.
This study presents a new data set of hourly PM2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM2.5. The resulting reanalysis from GSI provides an estimate of total PM2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R-2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set's fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM2.5 exposure.
Many fire weather index products have been developed to assist land managers, weather forecasters, and firefighters with anticipating weather conditions that may impact existing or potential new wildland fires in coming days. Most of these indices are designed to provide a single value for an entire 24-h period. Extreme wildfire activity in the western United States in recent years, including the impact of mesoscale and microscale phenomena such as thunderstorm gust frontal passages, radiative shading by dense smoke plumes, and pyrocumulonimbus development and collapse, as well as the advent of operational convection-allowing model forecasts, has highlighted the need for a more frequently updated index. In this study, we present a proof of concept for an hourly fire weather index developed specifically for application within a rapidly updating convection-allowing model. The index, termed the hourly wildfire potential (HWP), is developed based on observations of fire radiative power (FRP) from polar-orbiting satellites during large western U.S. wildfires in 2018 and 2020 and is evaluated against a merged dataset of FRP from polar-orbiting and geostationary satellites. The index is computed based on meteorological output from the NOAA operational High-Resolution Rapid Refresh (HRRR) model. The HWP index exhibits an improved representation of hourly FRP compared to a climatological approach and also shows promise for distinguishing between conditions associated with flaming and smoldering combustion. This work paves the way for improved prediction of wildfire smoke emissions in the coming hours and days.
A set of Surface Radiation Budget Network (SURFRAD) measurements across the lower 48 United States has allowed a closer inspection of weather model representations of downward shortwave radiation in the last several years. In this study, it is found that downward shortwave radiation (SW_) is excessive for the NOAA 3-km HRRR model at each of the 14 SURFRAD stations distributed across the lower United States when averaged over 2-month periods. Possible causes for this station-consistent SW_ bias error were hypothesized. Three were eliminated by this study and two were then evaluated in this study. We found that this error was not from clear-sky errors but from insufficient attenuation by clouds. It was also found that this cloud deficiency was partly caused by a dry bias in atmospheric water vapor initial conditions. New experiments using the hourly cycled HRRR model-assimilation system were designed and carried out for three seasons with modified data assimilation addressing the dry bias problem and reduction of effective radius for cloud water droplets for both explicit and subgrid-scale clouds. The assimilation and cloud optical parameter changes contributed similarly toward a combined reduced SW_ radiation bias by 80% in the fall season and 84% in the winter season but by only 35% in the summer season. Even with the improved data assimilation, a dry bias contributing to deficient clouds continues, which is a topic to be explored in a following study.
NOAA’s Global Systems Laboratory (GSL), in collaboration with other laboratories, is developing and testing a new high-resolution weather model known as the Rapid-Refresh Forecasting System (RRFS). This model, which uses the Finite Volume Cubed-Sphere Dynamical Core, features a grid covering all of North and Central America at 3 km horizontal resolution, with 65 vertical layers. The RRFS is initialized every hour through assimilation of the latest weather observations. It incorporates primary aerosol emissions from wildland fires and dust sources. The coupled RRFS-Smoke-Dust (RRFS-SD) model simulates 3D concentrations of smoke, fine and coarse dust aerosol species concurrently with the meteorology, and includes the aerosol radiative feedback. Hourly fire radiative power data from the Regional ABI and VIIRS fire Emissions (RAVE) product is ingested into RRFS to estimate biomass burning emissions and fire heat fluxes. Windblown dust emissions are parameterized by using the FENGSHA scheme. An experimental version of the RRFS-SD model is being tested by NOAA Environmental Modeling Center (EMC) in real time: https://rapidrefresh.noaa.gov/RRFS-SD/We will present an evaluation of the RRFS-SD model for several fire and dust case studies. Ground and aircraft-based in-situ and remote sensing data are extensively utilized to evaluate the model simulations of meteorology, smoke and dust fields. Additionally, we will present the radiative feedback of smoke and dust on the meteorological simulations in RRFS. The challenges and uncertainties affecting the smoke and dust forecasting will be discussed as well.
Probabilistic forecasts of excessive rainfall based on the fraction of High-Resolution Ensemble Forecast(HREF) members predicting precipitation above a given threshold are used widely in predicting excessive rainfall; how-ever, there is not yet a published study evaluating the skill of these forecasts. In this study, we document the performanceof these forecasts over a 3-yr period, including regional and seasonal variations in skill. Wefind that there is considerablesensitivity to how excessive rainfall events are defined, especially in regions with large differences in the number of exces-sive rainfall events between different datasets. When verifying against Stage IV exceedances offlashflood guidance (FFG),both the 0000 and 1200 UTC HREF probabilities of exceeding 6-h FFG thresholds exhibit a higher Brier skill score (BSS)than the operational 0900 UTC day-one excessive rainfall outlook (ERO) infive of eight regions in the contiguous UnitedStates (CONUS), while probabilities of exceedingfixed 6- or 12-h precipitation thresholds provide a higher BSS than theERO in another two regions. There is regional variability in the thresholds providing the highest BSS, with FFG (or 75%of FFG) generally providing the best forecasts in the eastern United States, butfixed thresholds providing the best fore-casts in the western United States. Only in the southeastern United States are threshold-based HREF forecasts unable tobeat the ERO. The 1200 UTC HREF-based forecasts using regionally optimal thresholds beat the ERO by 25%-30% interms of BSS. Our results suggest that HREF probabilities of exceeding precipitation thresholds have considerable valuefor excessive rainfall prediction.SIGNIFICANCE STATEMENT: Predicting excessive rainfall andflashflooding is a challenging problem. Opera-tional forecasters often use the fraction of high-resolution weather models predicting rainfall above a given threshold asone tool to guide their excessive rainfall outlooks, but it is unclear which threshold they should use or how much skillthe resulting probabilities have. Here, we evaluate these probabilities against observed excessive rainfall events andcompare them to operational forecasts. If we select the best-performing precipitation threshold in each region, wefindthat the model-based probabilities are more skillful than operational forecasts in seven of eight regions of the contigu-ous United States. These results inform forecasters about the best thresholds to use when developing their excessiverainfall outlooks
Background. NOAA's Hazard Mapping System (HMS) smoke product comprises smoke plumes digitised from satellite imagery. Recent studies have used HMS as a proxy for surface smoke presence. Aims We compare HMS with airport observations, air quality station measurements and model estimates of near-surface smoke. Methods. We quantify the agreement in numbers of smoke days and trends, regional discrepancies in levels of near-surface smoke fine particulate matter (PM2.5) within HMS polygons, and separation of total PM2.5 on smoke and non-smoke days across the contiguous US and Alaska from 2010 to 2021. Key results. We find large overestimates in HMS-derived smoke days and trends if we include light smoke plumes in the HMS smoke day definition. Outside the western US and Alaska, near-surface smoke PM2.5 within areas of HMS smoke plumes is low and almost indistinguishable across density categories, likely indicating frequent smoke aloft. Conclusions. Compared with airport, Environmental Protection Agency (EPA) and model-derived estimates, HMS most closely reflects surface smoke in the Pacific and Mountain regions and Alaska when smoke days are defined using only heavy plumes or both medium and heavy plumes. Implications We recommend careful consideration of biases in the HMS smoke product for air quality and public health assessments of fires.
Wildfires pose increasing risks to human health and properties in North America. Due to large uncertainties in fire emission, transport, and chemical transformation, it remains challenging to accurately predict air quality during wildfire events, hindering our collective capability to issue effective early warnings to protect public health and welfare. Here, we present a new real-time Hazardous Air Quality Ensemble System (HAQES) by leveraging various wildfire smoke forecasts from three U.S. federal agencies (NOAA, NASA, and Navy). Compared to individual models, the HAQES ensemble forecast significantly enhances forecast accuracy. To further enhance forecasting performance, a weighted ensemble forecast approach was introduced and tested. Compared to the unweighted ensemble mean, the multilinear regression weighted ensemble reduced fractional bias by 34% in the major fire regions, false alarm rate by 72%, and increased hit rate by 17%. Finally, we improved the weighted ensemble using quantile regression and weighted regression methods to enhance the forecast of extreme air quality events. The advanced weighted ensemble increased the PM2.5 exceedance hit rate by 55% compared to the ensemble mean. Our findings provide insights into the development of advanced ensemble forecast methods for wildfire air quality, offering a practical way to enhance decision -making support to protect public health. SIGNIFICANCE STATEMENT: Wildfires are a growing threat to health and safety in North America. Accurately predicting air quality during these events is crucial but challenging. In response, we have developed the real-time Hazardous Air Quality Ensemble System (HAQES), by combining forecasts from three U.S. federal agencies (NOAA, NASA, and Navy). HAQES significantly improves accuracy compared to individual models. Moreover, we further improve the wildfire air quality forecast by introducing the weighted ensemble method. The weighted ensemble reduced bias by 34% and false alarms by 72%, while increasing hit rates by 55%. HAQES advances our ability to protect public health during wildfire events.
Following the destructive Lahaina Fire in Hawaii, our team has modeled the wind and fire spread processes to understand the drivers of this devastating event. The results are in good agreement with observations recorded during the event. Extreme winds with high variability, a fire ignition close to the community, and construction characteristics led to continued fire spread in multiple directions. Our results suggest that available modeling capabilities can provide vital information to guide decision-making and emergency response management during wildfire events.
Flash flooding remains a challenging prediction problem, which is exacerbated by the lack of a universally accepted definition of the phenomenon. In this article, we extend prior analysis to examine the correspondence of various combinations of quantitative precipitation estimates (QPEs) and precipitation thresholds to observed occurrences of flash floods, additionally considering short-term quantitative precipitation forecasts from a convection-allowing model. Consistent with previous studies, there is large variability between QPE datasets in the frequency of "heavy" precipitation events. There is also large regional variability in the best thresholds for correspondence with reported flash floods. In general, flash flood guidance (FFG) exceedances provide the best correspondence with observed flash floods, although the best correspondence is often found for exceedances of ratios of FFG above or below unity. In the interior western United States, NOAA Atlas 14 derived recurrence interval thresholds (for the southwestern United States) and static thresholds (for the northern and central Rockies) provide better correspondence. The 6-h QPE provides better correspondence with observed flash floods than 1-h QPE in all regions except the West Coast and southwestern United States. Exceedances of precipitation thresholds in forecasts from the operational High-Resolution Rapid Refresh (HRRR) generally do not correspond with observed flash flood events as well as QPE datasets, but they outperform QPE datasets in some regions of complex terrain and sparse observational coverage such as the southwestern United States. These results can provide context for forecasters seeking to identify potential flash flood events based on QPE or forecast-based exceedances of precipitation thresholds.
Increasing impacts of wildfires on Western US air quality highlights the need for forecasts of smoke emissions based on dynamic modeled wildfires. This work utilizes knowledge of weather, fuels, topography, and firefighting, combined with machine learning and other statistical methods, to generate 1- and 2-day forecasts of fire radiative energy (FRE). The models are trained on data covering 2019 and 2021 and evaluated on data for 2020. For the 1-day (2-day) forecasts, the random forest model shows the most skill, explaining 48% (25%) of the variance in observed daily FRE when trained on all available predictors compared to the 2% (<0%) of variance explained by persistence for the extreme fire year of 2020. The random forest model also shows improved skill in forecasting day-to-day increases and decreases in FRE, with 28% (39%) of observed increase (decrease) days predicted, and increase (decrease) days are identified with 62% (60%) accuracy. Error in the random forest increases with FRE, and the random forest tends toward persistence under severe fire weather. Sensitivity analysis shows that near-surface weather and the latest observed FRE contribute the most to the skill of the model. When the random forest model was trained on subsets of the training data produced by agencies (e.g., the Canadian or US Forest Services), comparable if not better performance was achieved (1-day R-2 = 0.39-0.48, 2-day R-2 = 0.13-0.34). FRE is used to compute emissions, so these results demonstrate potential for improved fire emissions forecasts for air quality models.
Blowing snow is a hazard for motorists because it may rapidly reduce visibility. Numerical weather prediction models in the United States do not capture the movement of snow once it reaches the ground, but visibility reductions due to blowing snow can be diagnosed based on model-predicted land surface and environmental conditions that correlate with blowing snow occurrence. A recently developed diagnostic framework for forecasting blowing snow concentration and the associated visibility reduction is applied to High-Resolution Rapid Refresh (HRRR) and Rapid Refresh Forecast System (RRFS) model output including surface snow conditions to predict surface visibility reduction due to blowing snow. Twelve blowing snow events around Wyoming from 2018 to 2023 are examined. The analysis shows that visibility reductions due to blowing snow tend to be overpredicted, caused by the initial assumption of full driftability of the snowpack. This study refines the aging of the blowing snow reservoir with two methods. The first method estimates driftability based on time-varying snow density from the Rapid Update Cycle land surface model (RUC LSM) used in the HRRR and experimental RRFS models and is evaluated in a real-time context with the RRFS model. The second, complementary method diagnoses snowpack driftability using a process-based approach that requires data for recent snowfall, wind speed, and skin temperature. Compared to the full driftability assumption, this method shows limited improvements in forecasting skill. To improve model-based diagnosis of visibility reduction due to blowing snow, empirical work is needed to determine the relation between snowpack driftability and the recent history of snowfall and other weather conditions.
Background The record number of wildfires in the United States in recent years has led to an increased focus on developing tools to accurately forecast their impacts at high spatial and temporal resolutions. Aims The Warn-on-Forecast System for Smoke (WoFS-Smoke) was developed to improve these forecasts using wildfire properties retrieved from satellites to generate smoke plumes in the system. Methods The WoFS is a regional domain ensemble data assimilation and forecasting system built around the concept of creating short-term (0–6 h) forecasts of high impact weather. This work extends WoFS-Smoke by ingesting data from the GOES-16 satellite at 15-min intervals to sample the rapidly changing conditions associated with wildfires. Key results Comparison of experiments with and without GOES-16 data show that ingesting high temporal frequency data allows for wildfires to be initiated in the model earlier, leading to improved smoke forecasts during their early phases. Decreasing smoke plume intensity associated with weakening fires was also better forecast. Conclusions The results were consistent for a large fire near Boulder, Colorado and a multi-fire event in Texas, Oklahoma, and Arkansas, indicating a broad applicability of this system. Implications The development of WoFS-Smoke using geostationary satellite data allows for a significant advancement in smoke forecasting and its downstream impacts such as reductions in air quality, visibility, and potentially properties of severe convection.
The Marshall Fire on 30 December 2021 became the most destructive wildfire costwise in Colorado history as it evolved into a suburban firestorm in southeastern Boulder County, driven by strong winds and a snow-free and drought -influenced fuel state. The fire was driven by a strong downslope windstorm that maintained its intensity for nearly 11 hours. The southward movement of a large-scale jet axis across Boulder County brought a quick transition that day into a zone of upper-level descent, enhancing the midlevel inversion providing a favorable environment for an amplifying downstream mountain wave. In several aspects, this windstorm did not follow typical downslope windstorm behavior. NOAA rapidly updating numerical weather prediction guidance (including the High-Resolution Rapid Refresh) provided operationally useful forecasts of the windstorm, leading to the issuance of a High-Wind Warning (HWW) for eastern Boulder County. No Red Flag Warning was issued due to a too restrictive relative humidity criterion (already published alternatives are recommended); however, owing to the HWW, a countywide burn ban was issued for that day. Consideration of spatial (vertical and horizontal) and temporal (both valid time and initialization time) neighborhoods allows some quantification of forecast uncertainty from determin-istic forecasts-important in real-time use for forecasting and public warnings of extreme events. Essentially, dimensions of the deterministic model were used to roughly estimate an ensemble forecast. These dimensions including run-to-run consistency are also important for subsequent evaluation of forecasts for small-scale features such as downslope windstorms and the tropospheric features responsible for them, similar to forecasts of deep, moist convection and related severe weather.