Reliable power systems are essential to modern life, as severe storms continue to disrupt grid stability and cause widespread outages. Predicting storm outages enables utilities and emergency managers to pre-stage resources and improve resilience. However, the several days of forecast lead time typically needed for preparedness significantly affect the accuracy of outage predictions. This study investigates the impact of forecast lead time on the error propagation of a Gradient Boosting Machine (GBM)-based outage prediction model (OPM) driven by Weather Research and Forecasting (WRF) model forecasts and analysis predictions. We evaluate three error-analysis scenarios: FFAP (forecast vs. analysis-based outage predictions), FFAO (forecast vs. actual outages), and LFAO (leave-one-storm-out forecast vs. actual outages). Model performance is compared using Mean Absolute Percentage Error (MAPE) and Centered Root-Mean-Square Error (CRMSE) across short (12 h–1 d), medium (2–3 d), and long (4–5 d) forecast lead-time categories, with the long category representing the upper end of the medium-range forecast window relevant to operational preparedness. The results show that forecast lead time substantially affects outage prediction accuracy, but the magnitude depends on the evaluation setup. In the controlled FFAP scenario, CRMSE increased by approximately 110% as lead time increased, from 259 to 543 outages, isolating the effect of weather forecast degradation. In the more operational LFAO scenario, CRMSE was already high at short lead times, increasing from 847 to 920 outages, indicating that model generalization error dominates once storms are unseen. Across scenarios, LFAO errors were 51% higher than FFAO errors at short lead times, highlighting the importance of testing outage models under unseen-event conditions. These results quantify how forecast degradation and model generalization jointly shape the reliability of outage prediction and provide practical guidance for lead-time-aware storm preparedness.
Machine learning algorithms have shown promise in reducing bias in wind gust predictions, while still under-predicting high gusts. Uncertainty quantification (UQ) supports this issue by identifying when predictions are reliable or need cautious interpretation. Using data from 61 extratropical storms in the Northeastern USA, we introduce evidential neural network (ENN) as a novel approach for UQ in gust predictions, leveraging atmospheric variables from the Weather Research and Forecasting (WRF) model. Explainable AI techniques suggested that key predictive features contributed to higher uncertainty, which correlated strongly with storm intensity and spatial gust gradients. Compared to WRF, ENN demonstrated a 47 % reduction in RMSE and allowed the construction of gust prediction intervals without an ensemble, successfully capturing at least 95 % of observed gusts at 179 out of 266 stations. From an operational perspective, providing gust forecasts with quantified uncertainty enhances stakeholders' confidence in risk assessment and response planning for extreme gust events.
Statistical downscaling is a computationally efficient approach in producing a large ensemble of fine-resolution future climate projections to support local climate adaptation efforts. Here, we propose a deep learning-based framework, named BC_DLD, to downscale the Coupled Model Intercomparison Project phase 6 (CMIP6) Earth system model (ESM) outputs from the coarse native resolution to a locally relevant fine resolution at kilometer scale and apply it to the Northeast United States, a region experiencing rapid increase of heavy precipitation and in great need for locally actionable climate information. Employing a two-phase multivariate approach by utilizing an empirical-quantile delta mapping (E-QDM) approach for bias correction at a coarse resolution and deep convolutional neural networks (CNNs) for downscaling the corrected data to a finer spatial resolution, BC_DLD is designed to preserve the spatial characteristics from observations while retaining the temporal characteristics and climate change signals from ESMs. Here, we demonstrate the implementation of BC_DLD by downscaling the outputs of GFDL-ESM4 historical run and future shared socioeconomic pathway (SSP) 5-8.5 scenario over the Northeast, from -100 to -25 km [using fifth generation European Centre for Medium-Range Weather Forecasts atmospheric reanalysis (ERA5) as reference] and then further down to -6 km (using Livneh data as reference), and compare the downscaled data with an existing product, localized constructed analogs, version 2 (LOCA2). Despite using the same observational training data as LOCA2, BC_DLD produces more intense and rapid future increase of extreme precipitation than LOCA2, alleviating the underestimation of extremes common among statistical downscaling products. The emergent relationship between extreme precipitation intensity and temperature in BC_DLD more closely resembles the observations, further establishing its efficacy in producing useful and usable future climate projections.
Winter storms present significant hazards across the Northeast United States, often disrupting daily life. Numerical modeling of these storms is an important component for understanding the physical processes that cause significant impacts and for predicting their effects ahead of time. Initial and boundary conditions are an essential component to limited-area modeling; variability in these conditions can significantly alter the simulations. While previous modeling studies have investigated sensitivities in model physics, there has been limited exploration of the impact of different initial condition sources; differences within these sources can include horizontal and vertical resolution, data assimilation schemes, and domain. This study aims at identifying the sources of variability from the initialized atmospheric fields within four different sets of initial conditions and their impact on the prediction of winter precipitation processes. The key finding was that relative humidity across different initial and boundary conditions produced the most uncertainty on the model simulation, while variability in temperature or synoptic conditions had a minor role. To explain the precipitation differences seen during the simulations, vertical profiles of relative humidity and temperature were connected to microphysical hydrometeor species tracked within the model. The findings suggested that relative humidity differences are heavily linked to precipitation accumulation discrepancies and were the main source of variability from the initial conditions. These results call for the development of more accurate relative humidity profiles for model initial and boundary conditions.
Wind, often overlooked in climate change impact assessments, is now a crucial atmospheric variable due to our imminent reliance on renewable energy sources like wind energy. Therefore, establishing a wind speed database in higher temporal and spatial resolution is of great importance to assess wind power in historical periods and future projections. This study focuses on enhancing wind data representation and reliability of statistically downscaled model outputs by refining grid spacing from coarser to a desirable fi ner scale. Employing the Ensemble Generalized Analog Regression Downscaling (En-GARD) technique, we downscaled near-surface wind speed across a northeast (NE)U.S. domain that combines onshore and offshore wind farm areas. We tested multiple atmospheric variable combinations and identified near-surface wind speed (wind speed at 10 m), longitudinal and latitudinal wind at 10 m, and atmospheric temperature at 2 m, as the set that effectively guides the selection of analog days for the analog regression downscaling approach. Validation involved using ERA5 atmospheric reanalysis products (31 km) and high-resolution (4 km) Weather Research and Forecasting (WRF) Model simulations. Downscaling was conducted using a leave-one-year-out approach over 12 years (2001-12). Statistical analysis of the downscaled output over this period demonstrated a significant correlation and low error metrics (less than 20% for normalized RMSE and 16% for normalized MAE), indicating the reliability of the method and paving the way for its future application to downscale climate projection outputs. SIGNIFICANCE STATEMENT: Wind speed has become a significant atmospheric variable due to our society's imminent reliance on renewable energy resources like wind power. Often overlooked in climate change assessments, wind speed at a local scale over wind farm locations plays a crucial role in assessing and planning wind power resources in a future climate. We are providing the successful validation of a statistical downscaling methodology for near-surface wind speed at 4-km grid spacing using reanalysis data. The validity of the methodology gives confidence to downscale climate projection models and assess changes in wind speed for future climate scenarios in a follow-up study.
Reliable power forecasts are essential for the grid integration of offshore wind. This work presents a physics-based forecasting framework that couples mesoscale numerical weather prediction with large-eddy simulation (LES) and an actuator-disk turbine representation to predict farm-scale flows and power under realistic atmospheric conditions. Mean meteorological profiles from the Weather Research and Forecasting model drive a concurrent-precursor LES generating turbulent inflow consistent with the evolving boundary layer, while a main LES resolves turbulence and wake formation within the wind farm. The LES configuration and turbine-forcing implementation are validated against canonical single- and multi-turbine benchmarks, showing close agreement in wake deficits and recovery trends. The framework is then demonstrated for the South Fork Wind project (12 turbines, similar to 132 MW) using a set of time-varying cases over a 24 h period. Simulations reproduce hub-height wind variability, row-to-row power differences associated with wake interactions, and turbine-level power fluctuations (order 1 MW) that converge with appropriate averaging windows. The results illustrate how an LES-augmented hierarchical modeling system can complement conventional forecasting by providing physically interpretable flow fields and power estimates at operational scales.
Predicting snowfall accumulation in the Northeast United States, especially during extreme weather events like rapidly intensifying storms, is a significant challenge. Accurate snowfall predictions are crucial for public safety, infrastructure planning, and economic stability, yet they are difficult to achieve due to the complex processes of snowfall formation and forecasting. This study focused on the impact of enhancing temporal resolution and integrating multi-model ensemble data to improve snowfall prediction accuracy. We combined outputs from the Weather Research and Forecasting (WRF) model with machine learning (ML) algorithms and gridded snowfall products. Specifically, we explored the impact of finer temporal resolutions, such as 6-hour versus 24-hour feature intervals, on snowfall predictions and examined the impact of incorporating ensemble snowfall data that provided various percentile snowfall accumulations and probabilities of exceeding specific thresholds. Results demonstrated that the 6-hour model significantly reduced the overall prediction error, particularly during rapidly intensifying storms, by up to 30%. The inclusion of ensemble data further enhanced the prediction of 24hour snowfall, particularly in reducing the bias. Despite these advancements, challenges persist in accurately forecasting heavy snowfall amounts and capturing complex atmospheric dynamics during extreme events.
We integrate observations and simulated data from physics-based models with observations and machine learning (ML) algorithms to assess and predict lake dissolved oxygen (DO) and Apparent Oxygen Utilization (AOU). DO is a proxy of hypoxia, and AOU a proxy of respiration processes and biological activity. Weather, hydrology, and agroecosystem data were used to understand how various environmental drivers impact hypoxic conditions in Lake Erie for a 15-year period. We utilized outputs from various physics-based models and developed ML models with Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) to predict DO and AOU and rank the importance of influential explanatory variables. The ML models were able to accurately predict bottom DO and AOU, and identified thermal stratification as the most influential environmental variable, followed by mineralized phosphorus soil content. RF and XGBoost were not statistically different, therefore, we recommend the use of either ML algorithm to study hypoxic lake conditions.
We investigated the predictive capability of various configurations of the Weather Research and Forecasting (WRF) model version 4.4, to predict hub height offshore wind speed and wind power density in the Northeast US wind farm lease areas. The selected atmospheric conditions were high-pressure systems (anticyclones) coinciding with wind speed below the cut-in wind turbine threshold. There are many factors affecting the potential of offshore wind power generation, one of them being low winds, namely wind droughts, that have been present in future climate change scenarios. The efficiency of high-resolution hub height wind prediction for such events has not been extensively investigated, even though the anticipation of such events will be important in our increased reliance on wind and solar power resources in the near future. We used offshore wind observations from the Woods Hole Oceanographic Institution’s (WHOI) Air–Sea Interaction Tower (ASIT) located south of Martha’s Vineyard to assess the impact of the initial and boundary conditions, number of model vertical levels, and inclusion of high-resolution sea surface temperature (SST) fields. Our focus has been on the influence of the initial and boundary conditions (ICBCs), SST, and model vertical layers. Our findings showed that the ICBCs exhibited the strongest influence on hub height wind predictions above all other factors. The NAM/WRF and HRRR/WRF were able to capture the decreased wind speed, and there was no single configuration that systematically produced better results. However, when using the predicted wind speed to estimate the wind power density, the HRRR/WRF had statistically improved results, with lower errors than the NAM/WRF. Our work underscored that for predicting offshore wind resources, it is important to evaluate not only the WRF predictive wind speed, but also the connection of wind speed to wind power.
Accurate snowfall prediction is crucial for enhancing preparedness and resilience in the Northeast United States during winter weather events. This study introduces a novel approach that integrates Machine Learning (ML) models with atmospheric variables from the Weather Research and Forecasting (WRF) model to improve snowfall forecasts in the region. The significance lies in bridging the gap between physics-based Numerical Weather Prediction (NWP) models and the versatility of ML models, offering a promising advancement in winter storm predictions. Atmospheric variables related to 32 winter storms simulated by WRF, were used to feed Random Forest (RF) and Extreme Gradient Boosting (XGBoost) algorithms to predict 24 h accumulations of snowfall using the National Snowfall Analysis (NSA) product as reference. The comprehensive results revealed that the integrated approach provided more accurate snowfall prediction than the WRF and Air Force Weather Agency (AFWA) diagnostic, but faced challenges in the precision estimated by under-dispersion of the results. The WRF/XGBoost reduced the RMSE by 10.34 %, whereas WRF/RF reduced RMSE by 9.72 %, compared to WRF/AFWA. Similarly, WRF/XGBoost and WRF/RF increased the correlation coefficient by 12 % and 11 %. The most important variables in both ML algorithms were liquid water equivalent precipitation (LWE), snow ratio, wet-bulb temperature, temperature, and humidity at different heights, pressure tendency and wind speed, validating their ability to discern crucial factors influencing snowfall. Specific cases highlight the integrated approach's effectiveness in addressing challenges faced by traditional diagnostics, including LWE overestimation and warm temperature bias. However, limitations emerge in capturing storms characterized by anomalous atmospheric conditions and high snowfall values, most likely attributed to underrepresentation in the training data. By highlighting strengths and acknowledging limitations, this integrated approach holds the potential for improved snowfall prediction for future storms in the region compared to the traditional approach.
Wind gusts are often associated with severe hazards and can cause structural and environmental damages, making gust prediction a crucial element of weather forecasting services. In this study, we explored the utilization of machine learning (ML) algorithms integrated with numerical weather prediction outputs from the Weather Research and Forecasting (WRF) Model, to align the estimation of wind gust potential with observed gusts. We have used two ML algorithms, namely, random forest (RF) and extreme gradient boosting (XGB), along with two statistical techniques: generalized linear model with identity link function (GLM-Identity) and generalized linear model with log link function (GLM-Log), to predict storm wind gusts for the northeast (NE) United States. We used 61 simulated extratropical and tropical storms that occurred between 2005 and 2020 to develop and validate the ML and statistical models. To assess the ML model performance, we compared our results with postprocessed gust potential from WRF. Our findings showed that ML models, especially XGB, performed significantly better than statistical models and Unified Post Processor for the WRF (WRF-UPP) Model and were able to better align predicted with observed gusts across all storms. The ML models faced challenges capturing the upper tail of the gust distribution, and the learning curves suggested that XGB was more effective than RF in generating better predictions with fewer storms.
Tributary phosphorus (P) loads are one of the main drivers of eutrophication problems in freshwater lakes. Being able to predict P loads can aid in understanding subsequent load patterns and elucidate potential degraded water quality conditions in downstream surface waters. We demonstrate the development and performance of an integrated multimedia modeling system that uses machine learning (ML) to assess and predict monthly total P (TP) and dissolved reactive P (DRP) loads. Meteorological variables from the Weather Research and Forecasting (WRF) Model, hydrologic variables from the Variable Infiltration Capacity model, and agricultural management practice variables from the Environmental Policy Integrated Climate agroecosystem model are utilized to train the ML models to predict P loads. Our study presents a new modeling methodology using as testbeds the Maumee, Sandusky, Portage, and Raisin watersheds, which discharge into Lake Erie and contribute to significant P loads to the lake. Two models were built, one for TP loads using 10 environmental variables and one for DRP loads using nine environmental variables. Both models ranked streamflow as the most important predictive variable. In comparison with observations, TP and DRP loads were predicted very well temporally and spatially. Modeling results of TP loads are within the ranges of those obtained from other studies and on some occasions more accurate. Modeling results of DRP loads exceed performance measures from other studies. We explore the ability of both ML-based models to further improve as more data become available over time. This integrated multimedia approach is recommended for studying other freshwater systems and water quality variables using available decadal data from physics-based model simulations.
Tropical storm Isaias (2020) moved quickly northeast after its landfall in North Carolina and caused extensive damage to the east coast of the United States, with electric power distribution disruptions, infrastructure losses and significant economic and societal impacts. Improving the real-time prediction of tropical storms like Isaias can enable accurate disaster preparedness and strategy. We have explored the configuration, initialization and physics options of the Weather Research and Forecasting (WRF) model to improve the deterministic forecast for Isaias. The model performance has been evaluated based on the forecast of the storm track, intensity, wind and precipitation, with the support from in situ measurements and stage IV remote sensing products. Our results indicate that the Global Forecasting System (GFS) provides overall better initial and boundary conditions compared to the North American Model (NAM) for wind, mean sea level pressure and precipitation. The combination of tropical suite physics options and GFS initialization provided the best forecast improvement, with error reduction of 36% and an increase of the correlation by 11%. The choices for model spin-up time and forecast cycle did not affect the forecast of the storm significantly. In order to check the consistency of the result found from the investigation related to TS Isaias, Irene (2011), Henri (2021) and Elsa (2021), three other tropical storms, were also investigated. Similar to Isaias, these storms are simulated with NAM and GFS initialization and different physics options. The overall results for Henri and Elsa indicate that the models with GFS initialization and tropical suite physics reduced error by 44% and 57%, respectively, which resonates with the findings from the TS Isaias investigation. For Irene, the initialization used an older GFS version and showed increases in error, but applying the tropical physics option decreased the error by 20%. Our recommendation is to consider GFS for the initialization of the WRF model and the tropical physics suite in a future tropical storm forecast for the NE US.
High concentrations of aeolian dust affect the air quality and climate in large regions across Northern Africa, the Middle East, and parts of Asia. To assess the environmental impacts, numerical models have been developed that include mineral dust emissions, atmospheric transport and chemistry, and deposition processes. Since the dust can disperse across continents and oceans, there is a need to model a large geographical area. Here we present a state-of-the-art global atmospheric chemistry–climate model, with detailed representations of these processes. One unique model feature is the chemical interaction of dust with air pollution (chemical aging), which alters the microphysics of particles relevant for their atmospheric lifetime, e.g., the hygroscopic growth behavior, optical properties, and aerosol–cloud interactions, thus influencing the hydrologic cycle and climate. Based on recent developments and published results, we present a comparison of model calculations with satellite and ground-based remote sensing data as well as surface observations of dust concentrations and deposition. The model results are used to evaluate the consequences of aeolian dust for climate and public health.
Eutrophication and excessive algal growth pose a threat on aquatic organisms and the health of the public, environment, and the economy. Understanding what drives excessive algal growth can inform mitigation measures and aid in advance planning to minimize impacts. We demonstrate how simulated data from weather, hydrological, and agroecosystem numerical prediction models can be combined with machine learning (ML) to assess and predict Chlorophyll a (Chl a) concentrations, a proxy for lake eutrophication and algal biomass. The study area is Lake Erie for a 16-year period, 2002-2017. A total of 20 environmental variables from linked and coupled physical models are used as input features to train the ML model with Chl a observations from 16 measuring stations. Included are meteorological variables from the Weather Research and Forecasting (WRF) model, hydrological variables from the Variable Infiltration Capacity (VIC) model, and agricultural management practice variables from the Environmental Policy Integrated Climate (EPIC) agroecosystem model. The consolidation of these variables is conducive to a successful prediction of Chl a. Aside from the synergistic effects that weather, hydrology, and fertilizers have on eutrophication and excessive algal growth, we found that the application of different forms of both P and N fertilizers are highly ranked for the prediction of Chl a concentration. The developed ML model successfully predicts Chl a with a coefficient of determination of 0.81, bias of -0.12 μg/l and RMSE of 4.97 μg/l. The developed ML-based modeling approach can be used for impact assessment of agriculture practices in a changing climate that affect Chl a concentrations in Lake Erie.
Regional-scale air quality models are being used for studying the sources, composition, transport, transformation, and deposition of fine particulate matter (PM2.5). The availability of decadal air quality simulations provides a unique opportunity to explore sophisticated model evaluation techniques rather than relying solely on traditional operational evaluations. In this study, we propose a new approach for process-based model evaluation of speciated PM2.5 using improved complete ensemble empirical mode decomposition with adaptive noise (improved CEEMDAN) to assess how well version 5.0.2 of the coupled Weather Research and Forecasting model-Community Multiscale Air Quality model (WRF-CMAQ) simulates the time-dependent long-term trend and cyclical variations in daily average PM2.5 and its species, including sulfate (SO4), nitrate (NO3), ammonium (NH4), chloride (Cl), organic carbon (OC), and elemental carbon (EC). The utility of the proposed approach for model evaluation is demonstrated using PM2.5 data at three monitoring locations. At these locations, the model is generally more capable of simulating the rate of change in the long-term trend component than its absolute magnitude. Amplitudes of the sub-seasonal and annual cycles of total PM2.5, SO4, and OC are well reproduced. However, the time-dependent phase difference in the annual cycles for total PM2.5, OC, and EC reveals a phase shift of up to half a year, indicating the need for proper temporal allocation of emissions and for updating the treatment of organic aerosols compared to the model version used for this set of simulations. Evaluation of sub-seasonal and interannual variations indicates that CMAQ is more capable of replicating the sub-seasonal cycles than interannual variations in magnitude and phase.
Regional air quality models are widely being used to understand the spatial extent and magnitude of the ozone non-attainment problem and to design emission control strategies needed to comply with the relevant ozone standard through direct emission perturbations. In this study, we examine the manageable portion of ground-level ozone using two simulations of the Community Multiscale Air Quality (CMAQ) model for the year 2010 and a probabilistic analysis approach involving 29 years (1990-2018) of historical ozone observations. The modeling results reveal that the reduction in the peak ozone levels from total elimination of anthropogenic emissions within the model domain is around 13-21 ppb for the 90th-100th percentile range of the daily maximum 8-hr ozone concentrations across the contiguous United States (CONUS). Large reductions in the 4th highest 8-hr ozone are seen in the regions of West (interquartile range (IQR) of 17-33%), South (IQR 22-34%), Central (IQR 19-31%), Southeast (IQR 25-34%), and Northeast (IQR 24-37%). However, sites in the western portion of the domain generally show smaller reductions even when all anthropogenic emissions are removed, possibly due to the strong influence of global background ozone, including sources such as intercontinental ozone transport, stratospheric ozone intrusions, wildfires, and biogenic precursor emissions. Probabilistic estimates of the exceedances for several hypothetical thresholds of the 4th highest 8-hr ozone indicate that, in some areas, exceedances of such hypothetical thresholds may occur even with no anthropogenic emissions due to the ever-present atmospheric stochasticity and the current global tropospheric ozone burden. Implications: Because air pollution is intricately linked to adverse health effects, National Ambient Air Quality Standards (NAAQS) have been established for criteria pollutants to safeguard human health and the environment. Areas not in compliance with the relevant standards are required to develop plans and policies to reduce their air pollution levels. Regional-scale air quality models are currently being used routinely to inform policies to identify the emissions reduction required to meet and maintain the NAAQS throughout the country. This paper examines the feasibility of the 4th highest ozone, which is used to derive the ozone design value for NAAQS, complying with various current and hypothetical 8-hr ozone thresholds over CONUS based on the information embedded in 29 years of historical ozone observations and two modeling scenarios with and without anthropogenic emissions loading.
Regional-scale air quality models are being used fo r studying the sources, composition, transport, 10 transformation, and deposition of fine particulate matter (PM2.5). The availability of decadal air quality simulati ons 11 provides a unique opportunity to explore sophistica ted model evaluation techniques rather than relying solely on 12 traditional operational evaluations. In this study, we propose a new approach for process-based model evaluation of 13 speciated PM2.5 using improved Complete Ensemble Empirical Mode De composition with Adaptive Noise (improved 14 CEEMDAN) to assess how well version 5.0.2 of the co upled Weather Research and Forecasting model Comm unity 15 Multiscale Air Quality model (WRF-CMAQ) simulates t he time-dependent long-term trend and cyclical vari ations in 16 the daily average PM 2.5 and its species, including sulfate (SO 4), nitrate (NO3), ammonium (NH4), chloride (Cl) organic 17 carbon (OC) and elemental carbon (EC) . The utility of the proposed approach for model evaluation is d emonstrated 18 using PM2.5 data at three monitoring locations. At these locati ons, the model is generally more capable of simulat ing 19 the rate of change in the long-term trend component than its absolute magnitude. Amplitudes of the sub easonal and 20 annual cycles of total PM 2.5, SO4 and OC are well reproduced. However, the time-depe ndent phase difference in the 21 annual cycles for total PM 2.5, OC and EC reveal a phase shift of up to half year , indicating the need for proper temporal 22 allocation of emissions and for updating the treatm ent of organic aerosols compared to the model versi on used for this 23 set of simulations. Evaluation of sub-seasonal and inter-annual variations indicates that CMAQ is more capable of 24 replicating the sub-seasonal cycles than inter-annu l variations in magnitude and phase. 25
Regional-scale air pollution models are routinely being used world-wide for research, forecasting air quality, and regulatory purposes. It is well recognized that there are both reducible (systematic) and irreducible (unsystematic) errors in the meteorology-atmospheric chemistry modeling systems. The inherent (random) uncertainty stems from our inability to properly characterize stochastic variations in atmospheric dynamics and chemistry, and from the incommensurability associated with comparisons of the volume-averaged model estimates with point measurements. Because these stochastic variations are not being explicitly simulated in the current generation of regional-scale meteorology-air quality models, one should expect to find differences between the model estimates and corresponding observations. This paper presents an observation-based methodology to determine the expected errors from current generation regional air quality models even when the model design, physics, chemistry, and numerical analysis, as well as its input data, were "perfect". To this end, the short-term synoptic-scale fluctuations embedded in the daily maximum 8-hr ozone time series are separated from the longer-term forcing using a simple recursive moving average filter. The inherent uncertainty attributable to the stochastic nature of the atmosphere is determined based on 30+ years of historical ozone time series data measured at various monitoring sites in the contiguous United States. The results reveal that the expected root mean square error at the median and 95th percentile is about 2 ppb and 5 ppb, respectively, even for "perfect" air quality models driven with "perfect" input data. Quantitative estimation of the limit to the model's accuracy will help in objectively assessing the current state-of-the-science in regional air pollution models, measuring progress in their evolution, and providing meaningful and firm targets for improvements in their accuracy relative to ambient measurements.