Responding to the call for new verification methods in a recent editorial in Weather and Forecasting, this study proposed two new verification metrics to quantify the forecast challenges that a user faces in decision-making when using ensemble models. The measure of forecast challenge (MFC) combines forecast error and uncertainty information together into one single score. It consists of four elements: ensemble mean error, spread, nonlinearity, and outliers. The cross correlation among the four elements indicates that each element contains independent information. The relative contribution of each element to the MFC is analyzed by calculating the correlation between each element and MFC. The biggest contributor is the ensemble mean error, followed by the ensemble spread, nonlinearity, and outliers. By applying MFC to the predictability horizon diagram of a forecast ensemble, a predictability horizon diagram index (PHDX) is defined to quantify how the ensemble evolves at a specific location as an event approaches. The value of PHDX varies between 1.0 and −1.0. A positive PHDX indicates that the forecast challenge decreases as an event nears (type I), providing creditable forecast information to users. A negative PHDX value indicates that the forecast challenge increases as an event nears (type II), providing misleading information to users. A near-zero PHDX value indicates that the forecast challenge remains large as an event nears, providing largely uncertain information to users. Unlike current verification metrics that verify at a particular point in time, PHDX verifies a forecasting process through many forecasting cycles. Forecasting-process-oriented verification could be a new direction in model verification. The sample ensemble forecasts used in this study are produced from the NCEP global and regional ensembles.
Test beds have emerged as a critical mechanism linking weather research with forecasting operations. The U.S. Weather Research Program (USWRP) was formed in the 1990s to help identify key gaps in research related to major weather prediction problems and the role of observations and numerical models. This planning effort ultimately revealed the need for greater capacity and new approaches to improve the connectivity between the research and forecasting enterprise. Out of this developed the seeds for what is now termed test beds. While many individual projects, and even more broadly the NOAA/National Weather Service (NWS) Modernization, were successful in advancing weather prediction services, it was recognized that specific forecast problems warranted a more focused and elevated level of effort. The USWRP helped develop these concepts with science teams and provided seed funding for several of the test beds described. Based on the varying NOAA mission requirements for forecasting, differences in the organizational structure and methods used to provide those services, and differences in the state of the science related to those forecast challenges, test beds have taken on differing characteristics, strategies, and priorities. Current test bed efforts described have all emerged between 2000 and 2011 and focus on hurricanes (Joint Hurricane Testbed), precipitation (Hydrometeorology Testbed), satellite data assimilation (Joint Center for Satellite Data Assimilation), severe weather (Hazardous Weather Testbed), satellite data support for severe weather prediction (Short-Term Prediction Research and Transition Center), mesoscale modeling (Developmental Testbed Center), climate forecast products (Climate Testbed), testing and evaluation of satellite capabilities [Geostationary Operational Environmental Satellite-R Series (GOES-R) Proving Ground], aviation applications (Aviation Weather Testbed), and observing system experiments (OSSE Testbed).
The NOAA Hazardous Weather Testbed (HWT) conducts annual spring forecasting experiments organized by the Storm Prediction Center and National Severe Storms Laboratory to test and evaluate emerging scientific concepts and technologies for improved analysis and prediction of hazardous mesoscale weather. A primary goal is to accelerate the transfer of promising new scientific concepts and tools from research to operations through the use of intensive real-time experimental forecasting and evaluation activities conducted during the spring and early summer convective storm period. The 2010 NOAA/HWT Spring Forecasting Experiment (SE2010), conducted 17 May through 18 June, had a broad focus, with emphases on heavy rainfall and aviation weather, through collaboration with the Hydrometeorological Prediction Center (HPC) and the Aviation Weather Center (AWC), respectively. In addition, using the computing resources of the National Institute for Computational Sciences at the University of Tennessee, the Center for Analysis and Prediction of Storms at the University of Oklahoma provided unprecedented real-time conterminous United States (CONUS) forecasts from a multimodel Storm-Scale Ensemble Forecast (SSEF) system with 4-km grid spacing and 26 members and from a 1-km grid spacing configuration of the Weather Research and Forecasting model. Several other organizations provided additional experimental high-resolution model output. This article summarizes the activities, insights, and preliminary findings from SE2010, emphasizing the use of the SSEF system and the successful collaboration with the HPC and AWC. A supplement to this article is available online (DOI:10.1175/BAMS-D-11-00040.2)
A new strategy for generating and presenting model diagnostic fields from convection-allowing forecast models is introduced. The fields are produced by computing temporal-maximum values for selected diagnostics at each horizontal grid point between scheduled output times. The two-dimensional arrays containing these maximum values are saved at the scheduled output times. The additional fields have minimal impacts on the size of the output files and the computation of most diagnostic quantities can be done very efficiently during integration of the Weather Research and Forecasting Model. Results show that these unique output fields facilitate the examination of features associated with convective storms, which can change dramatically within typical output intervals of 1-3 h.
The impacts of assimilating radar data and other mesoscale observations in real-time, convection-allowing model forecasts were evaluated during the spring seasons of 2008 and 2009 as part of the Hazardous Weather Test Bed Spring Experiment activities. In tests of a prototype continental U. S.-scale forecast system, focusing primarily on regions with active deep convection at the initial time, assimilation of these observations had a positive impact. Daily interrogation of output by teams of modelers, forecasters, and verification experts provided additional insights into the value-added characteristics of the unique assimilation forecasts. This evaluation revealed that the positive effects of the assimilation were greatest during the first 3-6 h of each forecast, appeared to be most pronounced with larger convective systems, and may have been related to a phase lag that sometimes developed when the convective-scale information was not assimilated. These preliminary results are currently being evaluated further using advanced objective verification techniques.
The impact of radar-data assimilation in real-time, convection-allowing model forecasts was evaluated during the spring seasons of 2008 and 2009 as part of Hazardous Weather Testbed (HWT) Spring Experiment activities. Preliminary results suggest that an early prototype version of a CONUS-scale assimilation system had a positive impact on model forecasts, especially when organized convective activity was ongoing at the initial time. Daily interrogation of output by teams of modelers, forecasters, and verification experts provided additional insight into the value-added characteristics of the radar-assimilation forecasts. This evaluation revealed that the positive effect of the assimilation was greatest during the first 3-6 h of each forecast, appeared to be most pronounced with larger convective systems, and may have been related to a phase lag that sometimes developed when the convective-scale information was not assimilated. These preliminary results are currently being evaluated further using advanced objective verification techniques.
During the 2007 NOAA Hazardous Weather Testbed (HWT) Spring Experiment, the Center for Analysis and Prediction of Storms (CAPS) at the University of Oklahoma produced convection-allowing forecasts from a single deterministic 2-km model and a 10-member 4-km-resolution ensemble. In this study, the 2-km deterministic output was compared with forecasts from the 4-km ensemble control member. Other than the difference in horizontal resolution, the two sets of forecasts featured identical Advanced Research Weather Research and Forecasting model (ARW-WRF) configurations, including vertical resolution, forecast domain, initial and lateral boundary conditions, and physical parameterizations. Therefore, forecast disparities were attributed solely to differences in horizontal grid spacing. This study is a follow-up to similar work that was based on results from the 2005 Spring Experiment. Unlike the 2005 experiment, however, model configurations were more rigorously controlled in the present study, providing a more robust dataset and a cleaner isolation of the dependence on horizontal resolution. Additionally, in this study, the 2- and 4-km outputs were compared with 12-km forecasts from the North American Mesoscale (NAM) model. Model forecasts were analyzed using objective verification of mean hourly precipitation and visual comparison of individual events, primarily during the 21- to 33-h forecast period to examine the utility of the models as next-day guidance. On average, both the 2- and 4-km model forecasts showed substantial improvement over the 12-km NAM. However, although the 2-km forecasts produced more-detailed structures on the smallest resolvable scales, the patterns of convective initiation, evolution, and organization were remarkably similar to the 4-km output. Moreover, on average, metrics such as equitable threat score, frequency bias, and fractions skill score revealed no statistical improvement of the 2-km forecasts compared to the 4-km forecasts. These results, based on the 2007 dataset, corroborate previous findings, suggesting that decreasing horizontal grid spacing from 4 to 2 km provides little added value as next-day guidance for severe convective storm and heavy rain forecasters in the United States.
During the 2007 NOAA Hazardous Weather Testbed Spring Experiment, the Center for Analysis and Prediction of Storms (CAPS) at the University of Oklahoma produced a daily 10-member 4-km horizontal resolution ensemble forecast covering approximately three-fourths of the continental United States. Each member used the Advanced Research version of the Weather Research and Forecasting (WRF-ARW) model core, which was initialized at 2100 UTC, ran for 33 h, and resolved convection explicitly. Different initial condition (IC), lateral boundary condition (LBC), and physics perturbations were introduced in 4 of the 10 ensemble members, while the remaining 6 members used identical ICs and LBCs, differing only in terms of microphysics (MP) and planetary boundary layer (PBL) parameterizations. This study focuses on precipitation forecasts from the ensemble.The ensemble forecasts reveal WRF-ARW sensitivity to MP and PBL schemes. For example, over the 7-week experiment, the Mellor-Yamada-Janjic PBL and Ferrier MP parameterizations were associated with relatively high precipitation totals, while members configured with the Thompson MP or Yonsei University PBL scheme produced comparatively less precipitation. Additionally, different approaches for generating probabilistic ensemble guidance are explored. Specifically, a ``neighborhood'' approach is described and shown to considerably enhance the skill of probabilistic forecasts for precipitation when combined with a traditional technique of producing ensemble probability fields.
1. Introduction Throughout the history of numerical weather prediction (NWP), computer resources have increased to enable NWP models to run at progressively higher resolutions over increasingly large domains. using convection-allowing [no convective parameterization (CP)] configurations of the Weather Research and Forecasting (WRF) model with horizontal grid spacings of ~ 4 km have demonstrated the added value of these high-resolution models as forecast guidance tools for the prediction of heavy precipitation. Additionally, these experiments have revealed minimal adverse effects from running the WRF model at 4 km without CP, even though this grid spacing is too coarse to fully capture convective scale circulations. Given the success of these convection-allowing WRF forecasts, ~ 4 km convection-allowing models have become operational at the United States National Centers for Environmental Prediction (NCEP) in the form of " high-resolution window " deterministic forecasts produced by the Environmental Modeling Center (EMC) of NCEP.
Abstract The Super Outbreak of tornadoes over the central and eastern United States on 3–4 April 1974 remains the most outstanding severe convective weather episode on record in the continental United States. The outbreak far surpassed previous and succeeding events in severity, longevity, and extent. In this paper, surface, upper-air, radar, and satellite data are used to provide an updated synoptic and subsynoptic overview of the event. Emphasis is placed on identifying the major factors that contributed to the development of the three main convective bands associated with the outbreak, and on identifying the conditions that may have contributed to the outstanding number of intense and long-lasting tornadoes. Selected output from a 29-km, 50-layer version of the Eta forecast model, a version similar to that available operationally in the mid-1990s, also is presented to help depict the evolution of thermodynamic stability during the event.
Jason J. Levit*, Gregory W. Carbin, David R. Bright, John S. Kain, Steven J. Weiss, Russell S. Schneider, Michael C. Coniglio, Ming Xue, Kevin W. Thomas, Matthew E. Pyle, Morris L. Weisman NOAA/NWS Storm Prediction Center NOAA National Severe Storms Laboratory Center for Analysis and Prediction of Storms and University of Oklahoma NOAA/NWS Environmental Modeling Center National Center for Atmospheric Research
During the 2005 NOAA Hazardous Weather Testbed Spring Experiment two different high-resolution configurations of the Weather Research and Forecasting-Advanced Research WRF (WRF-ARW) model were used to produce 30-h forecasts 5 days a week for a total of 7 weeks. These configurations used the same physical parameterizations and the same input dataset for the initial and boundary conditions, differing primarily in their spatial resolution. The first set of runs used 4-km horizontal grid spacing with 35 vertical levels while the second used 2-km grid spacing and 51 vertical levels. Output from these daily forecasts is analyzed to assess the numerical forecast sensitivity to spatial resolution in the upper end of the convection-allowing range of grid spacing. The focus is on the central United States and the time period 18–30 h after model initialization. The analysis is based on a combination of visual comparison, systematic subjective verification conducted during the Spring Experiment, and objective metrics based largely on the mean diurnal cycle of the simulated reflectivity and precipitation fields. Additional insight is gained by examining the size distributions of the individual reflectivity and precipitation entities, and by comparing forecasts of mesocyclone occurrence in the two sets of forecasts. In general, the 2-km forecasts provide more detailed presentations of convective activity, but there appears to be little, if any, forecast skill on the scales where the added details emerge. On the scales where both model configurations show higher levels of skill—the scale of mesoscale convective features—the numerical forecasts appear to provide comparable utility as guidance for severe weather forecasters. These results suggest that, for the geographical, phenomenological, and temporal parameters of this study, any added value provided by decreasing the grid increment from 4 to 2 km (with commensurate adjustments to the vertical resolution) may not be worth the considerable increases in computational expense.
Valliappa Lakshmanan合作论文数University of Oklahoma2