When providing pollen forecasts to the community, there is a need to verify the accuracy of curated forecasts, but evaluation is not routinely reported. This study of the AusPollen Partnership compared multi-category grass pollen forecasts for up to six days ahead with daily airborne grass pollen concentrations measured in Brisbane, Canberra, Melbourne, and Sydney, Australia during four pollen seasons from 2016 to 2020. The accuracy of categorical grass pollen forecasts predicting grass pollen concentrations in the high and greater, or moderate and greater categories, were assessed as often applied in meteorology using Gerrity scores, equitable threat scores, false alarm ratios, success ratios, and probability of detection of correct category. The skill of grass pollen forecasts curated by aerobiologists were compared with two retrospectively calculated naïve reference forecast methods; climatology and persistence. For Brisbane and Melbourne, high or greater grass pollen levels occurred on average 32% and 22% of days, whereas for Canberra and Sydney, there were few high days, but moderate or greater pollen levels occurred on average 26% and 19% of days, respectively. Average annual Gerrity scores for curated forecasts of high or greater improved with experience from 0.20 to 0.66 in Brisbane, and from 0.39 to 0.55 in Melbourne between 2016 and 2019. Average Gerrity Scores for moderate or greater categories in Sydney were 0.45 and 0.43 in 2016 and 2018 respectively, and in Canberra were 0.34 and 0.41, in the same years. The skill of curated forecasts was usually better than persistence forecasts, but the accuracy of the curated forecasts decreased with longer lead times. Although persistence grass pollen forecasts consistently performed better than climatologies, persistence depends on previous day pollen concentrations being available. Short-term curated daily grass pollen forecasts of the AusPollen Partnership offer useful information for people with allergic rhinitis and asthma, to help facilitate behavioural change and reduce the health burden. There is a need in Australia to extend local pollen records through sustained pollen monitoring to track climate-related changes as well as improve reliability of daily pollen forecasts. Globally, continued evaluation will enable reporting of accurate pollen forecasts to community, clinicians and government stakeholders.
As models are refined and developed, it is imperative to have objective ways to evaluate forecast quality. This is important for comparing different model configurations or tracking performance over time. Greater computing power has allowed finer grid spacing and more explicit handling of previously unresolved circulations (which is crucial for distinguishing high impact events with intense peaks in wind or precipitation). The problem is, as grid spacing decreases, the traditional verification methods become swamped by small-scale errors and they often cannot discriminate between a somewhat-useful forecast and a totally useless forecast. For example, a high-resolution forecasted precipitation field may look very good and be quite useful, but if it is slightly offset from the observations, the traditional verification scores (such as critical success index and equitable threat score) will be dominated by false-alarms and misses due to slight displacement errors. Forecasted and observed events are unlikely to be matched up exactly on a point-bypoint basis and the forecast is “doubly-penalized” for false alarms and misses associated with what is essentially the same entity. We would like our verification metric to be sensitive to displacement errors, but at the same time not give an inordinate amount of weight to trivial deviations from the observations (truth). It is with this in mind that we look at several innovative approaches to spatial forecast verification. These methods, which were discussed at a verification workshop in Feb. 2007, can be divided into three broad categories: feature-based, neighborhood approach, and scale decomposition.
Statistical and case study - oriented comparisons of the quantitative precipitation nowcasting (QPN) schemes demonstrated during the first World Weather Research Programme (WWRP) Forecast Demonstration Project (FDP), held in Sydney, Australia, during 2000, served to confirm many of the earlier reported findings regarding QPN algorithm design and performance. With a few notable exceptions, nowcasting algorithms based upon the linear extrapolation of observed precipitation motion ( Lagrangian persistence) were generally superior to more sophisticated, nonlinear nowcasting methods. Centroid trackers [ Thunderstorm Identification, Tracking, Analysis and Nowcasting System ( TITAN)] and pattern matching extrapolators using multiple vectors ( Auto-nowcaster and Nimrod) were most reliable in convective scenarios. During widespread, stratiform rain events, the pattern matching extrapolators were superior to centroid trackers and wind advection techniques ( Gandolf, Nimrod).There is some limited case study and statistical evidence from the FDP to support the use of more sophisticated, nonlinear QPN algorithms. In a companion paper in this issue, Wilson et al. demonstrate the advantages of combining linear extrapolation with algorithms designed to predict convective initiation, growth, and decay in the Auto- nowcaster. Ebert et al. show that the application of a nonlinear scheme [ Spectral Prognosis ( S- PROG)] designed to smooth precipitation features at a rate consistent with their observed temporal persistence tends to produce a nowcast that is superior to Lagrangian persistence in terms of rms error. However, the value of this approach in severe weather forecasting is called into question due to the rapid smoothing of high-intensity precipitation features.
The Sydney 2000 Olympic Games World Weather Research Programme Forecast Demonstration Project (WWRP FDP) aimed to demonstrate the utility and impact of modern nowcast systems. The project focused on the use of radar processing systems and products for nowcasting, including severe weather. The forecast problems facing the Australian Bureau of Meteorology (BoM) on these short timescales during the FDP are briefly described. The observing system is then discussed and enhancements to the network that supported the Olympic Games forecast requirements and the WWRP FDP project are outlined. In particular, issues related to radar calibration and quality control are discussed in some detail. The paper concludes with a brief discussion on the observing system requirements to meet such modern nowcast systems, areas of further development, and impacts that the FDP had on BoM nowcasting systems. The need for end-to-end design of systems from data gathering, to analysis and product generation is emphasized.
The first World Weather Research Programme (WWRP) Forecast Demonstration Project (FDP), with a focus on nowcasting, was conducted in Sydney, Australia, from 4 September to 21 November 2000 during a period associated with the Sydney 2000 Olympic Games. Through international collaboration, nine nowcasting systems from the United States, United Kingdom, Canada, and Australia were deployed at the Sydney Office of the Bureau of Meteorology (BOM) to demonstrate the capability of modern forecast systems and to quantify the associated benefits in the delivery of a real-time nowcast service. On-going verification and impact studies supported by international committees assisted by the WWRP formed an integral part of this project. A description is given of the project, including component systems, the weather, and initial outcomes. Initial results show that the nowcasting systems tested were transferable and able to provide valuable information enhancing BOM nowcasts. The project provided for unprecedented interchange of concepts and ideas between forecasters, researchers, and end users in an operational framework where they all faced common issues relevant to real time nowcast decision making. A training workshop sponsored by the World Meteorological Organization (WMO) was also held in conjunction with the project so that other member nations could benefit from the FDP.
This paper investigates the use of an Artificial Neural Network (ANN) to estimate the six hour rainfall over the south east coast of Tasmania. ANN's are becoming increasingly prominent in many areas of weather forecasting due to their potential to capture the complex relationships between the many factors that contribute to certain weather conditions. The estimations produced by the ANN's were compared to one estimation technique and one forecast technique used by the Bureau of Meteorology. The results confirm that ANN's have the potential for successful application to the problem of rainfall estimation.
Measurement of polar cloud cover is important because of its strong radiative influence on the energy balance of the snow and ice surface. Conventional satellite cloud detection schemes often fail in the polar regions because the visible and thermal contrasts between cloud and surface are typically small. Nevertheless, experts looking at satellite imagery can distinguish clouds from the surface by examining the textural characteristics of the scene. This paper describes an automated pattern recognition algorithm winch identities regions of various surface and cloud types at high latitudes from visible, near-infrared, and infrared AVHRR satellite data. Five spectral features give information about the magnitude of albedos and brightness temperatures, while three textural features describe the variability and "bumpiness" in a scene. The maximum likelihood decision rule is used to classify that region into one of seven surface categories or 11 cloud categories. The algorithm was able to classify 870 training samples with a skill of 84%. Eighteen hundred artificer samples created using a Monte Carlo technique were classified with a skill of 92%, which represents the theoretical limit of class separability using the given features. Both the near-infrared information and the textural information proved to be especially useful in detecting high-latitude cloudiness. The algorithm experienced some difficulty identifying thin stratus over snow and ice and thin cirrus over land and water, situations which also prove difficult for most other cloud detection schemes. When tested on AVHRR imagery from a different date, the algorithm showed a skill of 83% as verified against the analyses of three independent experts. Significant variability was encountered among the experts, underlining the need for an objective routine. This algorithm performed more accurately thin others constructed with alternate feature sets corresponding to various existing cloud detection schemes.