ABSTRACT Guangxi, located along China's southern coast, is prone to typhoons and features complex terrain, making wind speed forecasting challenging. Accurate prediction of near‐surface maximum wind speed is crucial for improving wind energy utilization and supporting carbon neutrality goals. This study proposes a novel prediction model using the eXtreme Gradient Boosting (XGBoost) algorithm integrated with a Bayesian Optimization Algorithm (BOA) and based on the k‐nearest neighbor mutual information feature selection algorithm (KNN‐MIFSA). Data from 93 meteorological stations in Guangxi (2016–2021) with a 3‐h temporal resolution were used. The model incorporates dynamic and thermal factors, including high‐altitude and surface variables, to predict maximum wind speed. Two key improvements were made in the prediction modeling: (1) KNN‐MIFSA was employed to select highly correlated features and eliminate redundant variables, and (2) BOA was used to optimize XGBoost parameters, enhancing model generalizability. The improved model was tested for 6 prediction lead times (12–72 h) from 2020 to 2021. Results show that, after adjusting parameters and processing factors, the new model reduced the mean absolute error (MAE) by 18.9%–30.06% and the root mean square error (RMSE) by 40.18%–65.83% compared to the original XGBoost model. For maximum wind speeds above level 6, MAE and RMSE of the new model were reduced by up to 40.41% and 30.92%, respectively, across lead times (12–72 h). The model demonstrates consistent performance and significantly improved accuracy, offering a promising approach for wind speed prediction in regions with complex terrain.
This study focused on predicting the near-surface maximum wind speed using the eXtreme Gradient Boosting (XGBoost) model based on k-nearest neighbor mutual information feature selection. The data from 93 meteorological stations in Guangxi Province, with a temporal resolution of 3 h, were used for the prediction. By examining the effects of various dynamic and thermal factors, such as high altitudes and surface variables, on the prediction of maximum wind speed, a novel XGBoost-based prediction model for maximum wind speed was proposed. The model incorporates the k-nearest neighbor mutual information feature selection algorithm to choose the most relevant factors for accurate wind speed prediction. In the design of the prediction model, there are two main areas of improvement. First, a stepwise variable selection algorithm based on k-nearest neighbor mutual information estimation was employed, which selects relevant variables and removes weakly relevant variables through two steps, effectively eliminating redundant prediction characteristics that affect accuracy by screening the primary predictors and retaining important forecasting factors. Second, the Bayesian optimization algorithm was used to optimize the parameters in the XGBoost model, significantly enhancing the model's generalizability. The optimized and improved prediction model was utilized to model and research the near-surface maximum wind speed for 6 forecast lead times (12-72 h) at 93 meteorological stations. Comparative results of various forecast experiments using independent prediction samples from 2020 to 2021 demonstrated that the new model reduced the average mean absolute error (MAE) evaluation metric by 18.9% to 30.06% for the prediction results of the 93 stations. The root mean square error (RMSE) metric decreased by 40.18% to 65.83%. For the prediction of maximum wind speeds exceeding level 6, the MAE was reduced by 40.41%, 25.93%, 19.96%, 21.39%, 12.39%, and 8.55% for the 6 forecast lead times, respectively. The RMSE evaluation metric also decreased by 30.92%, 18.67%, 12.29%, 12.21%, 7.92%, and 2.39% for the respective lead times.
To enhance the forecasting capability for daily extreme wind speeds, particularly for winds exceeding force 8, this paper uses the “past 3 h gust” wind speed forecast output from the European Centre for Medium-Range Weather Forecasts (ECMWF) model as the primary input factor. Additionally, the paper addresses the extremely uneven sample distribution in the daily extreme wind speed series (samples with wind force above level 8 constitute a very small proportion of the total sample, while samples with wind force below level 5 constitute the vast majority). Moreover, the ECMWF model’s “past 3 h gust” wind speed forecast tends to overestimate low-level winds and underestimate high-level winds. Therefore, the paper leverages nearly five years of surface observations and ECMWF model “past 3 h gust” forecast data to develop a Tabnet-based daily extreme wind classification correction forecast model. The model’s input design includes previous observations, geographic information of the stations, ECMWF forecast fields, and previous forecast error terms. In the evaluation of an independent sample over one and a half years, the new correction forecast model reduces the mean absolute error (MAE) by 45.2% and the root mean square error (RMSE) by 25.7% compared to the interpolated ECMWF model. Furthermore, for wind force levels 1-5 and above 8-9, the new correction forecast model significantly improves the forecasting accuracy compared to the method using interpolated ECMWF forecast fields, demonstrating the feasibility of this forecasting approach.The model is constructed with a focus on overcoming the inherent limitations of the ECMWF model’s wind speed forecasts. By incorporating comprehensive input factors such as historical observation data, the geographical context of observation stations, and systematic forecast error corrections, the model aims to provide a more accurate prediction of extreme wind events. The primary challenge addressed by the model is the skewed distribution of wind force levels in the dataset, where extreme wind events are underrepresented. The innovative use of the Tabnet algorithm allows for a sophisticated analysis and adjustment of the forecast data, thus ensuring higher accuracy in predicting both low and high wind force levels. The independent validation over an extensive period highlights the robustness of the model. The significant reduction in MAE and RMSE underscores the model’s enhanced performance. Specifically, the accuracy improvements for the critical wind force levels 1-5 and 8-9 plus indicate the model’s practical applicability in real-world scenarios. This advancement is crucial for sectors reliant on precise wind forecasts, such as maritime operations, aviation, and disaster preparedness. The results clearly suggest that integrating historical data and addressing the ECMWF model’s biases can lead to substantial improvements in extreme wind speed forecasting. In conclusion, the development of the Tabnet-based correction forecast model represents a significant step forward in meteorological forecasting. By effectively addressing the biases and limitations of existing models, this new approach offers a more reliable tool for predicting extreme wind events.
Aiming at the lack of nonlinear intelligent computational modelling methods for the fixed-point and quantitative forecasting of typhoon gales in the current numerical forecast products, the paper takes the daily extreme winds at five typical representative meteorological stations (Guilin, Wuzhou, Longzhou, Nanning, Yulin) as the forecast object, and carries out the construction of a daily extreme wind forecast model based on multivariate linear regression (MR), support vector machine (SVM), fuzzy neural network (FNN), and the ground-based observation and reanalysis of the data of the typhoon in the past 40 years during the typhoon impact in Guangxi. The construction of the daily maximum wind prediction model based on multiple linear regression (MR), support vector machine (SVM) and fuzzy neural network (FNN) is carried out. The test results of the independent samples show that the FNN model has the smallest mean absolute error for the four stations of Guilin, Wuzhou, Longzhou, and Yulin in terms of the mean absolute error of the full-sample wind speed forecast, and the best overall forecast accuracy, while the MR forecast model has a better forecast capability for Nanning station, and the SVM model has an overall bias in the forecast effect, in which the FNN forecast model has a 1% to 1% reduction in mean absolute error compared with that of MR. The mean absolute errors of the FNN forecast model are reduced by 1%-29% (except for Nanning station); the mean absolute errors of the FNN forecast model are reduced by 6%-29% compared with the SVM forecast model. the mean absolute errors of the MR forecast model are reduced by 5%-13% compared with the SVM forecast model (except for Guilin station). The statistical results of the four evaluation indexes, including TS score, hit rate, null rate and forecast bias, for winds of magnitude 6 or above show that the FNN model has the highest and relatively stable prediction accuracy, followed by the MR scheme, and the SVM has the worst prediction effect among the three schemes. The fuzzy neural network has certain applicability to the prediction of very high wind speed, which can be a good reference for the prediction of daily very high wind speed on the ground during typhoons in Guangxi, and can provide theoretical references and empirical evidence basis for the later development of the research on the prediction of high wind disasters in Guangxi.
In the current short-term climate prediction of monthly precipitation, there is a lack of nonlinear data mining techniques and objective ensemble forecasting methods of machine learning. A new nonlinear deep learning ensemble objective forecasting model has been established by generating multiple long short-term memory neural networks (LSTMs) with the same expected output as the individual forecasters, and using the cooperative game Shapley value method to determine the weight coefficients of each forecaster in the ensemble forecasting. The forecasting modeling of the monthly precipitation forecasting model has been studied based on the July precipitation samples of 81 meteorological stations in Guangxi from 1960 to 2023, and using height fields, temperature fields, and sea surface temperature field as the basic forecasting factors for monthly precipitation. The experimental results show that under the same forecast modeling samples and forecast factor conditions, the newly established prediction model has higher predictive ability than linear stepwise regression prediction methods and a single LSTM model, demonstrating its applicability to nonlinear monthly precipitation prediction problems. Further analysis reveals that the introduction of storage unit states and gate structures in the hidden layer of the LSTM model enables the network to retain long-term states, making it more suitable for handling and predicting important problems with relatively long intervals and delays in time series. And the Shapley value method can improve the predictive ability of ensemble individuals and enhance the population diversity of ensemble individuals, thereby improving the predictive accuracy of ensemble forecasting models. Therefore, the generalization ability of this deep learning ensemble forecasting model is significantly improved, and the improvement of its forecasting ability has a reasonable analytical basis. There is no overfitting phenomenon in the practical short-term climate prediction business application of general neural network methods, and it has good practical application value.
This study focused on predicting the near-surface maximum wind speed using the eXtreme Gradient Boosting (XGBoost) model based on k-nearest neighbor mutual information feature selection. The data from 93 meteorological stations in Guangxi Province from 2016 to 2021, with a temporal resolution of 3 h, were used for the prediction. By examining the effects of various dynamic and thermal factors, such as high altitudes and surface variables, on the prediction of maximum wind speed, a novel XGBoost-based prediction model for maximum wind speed was proposed. The model incorporates the k-nearest neighbor mutual information feature selection algorithm to choose the most relevant factors for accurate wind speed prediction. In the design of the prediction model, there are two main areas of improvement. First, a stepwise variable selection algorithm based on k-nearest neighbor mutual information estimation was employed, which selects relevant variables and removes weakly relevant variables through two steps, effectively eliminating redundant prediction characteristics that affect accuracy by screening the primary predictors and retaining important forecasting factors. Second, the Bayesian optimization algorithm was used to optimize the parameters in the XGBoost model, significantly enhancing the model's generalizability. The optimized and improved prediction model was utilized to model and research the near-surface maximum wind speed for 6 forecast lead times (12–72 h) at 93 meteorological stations. Comparative results of various forecast experiments using independent prediction samples from 2020 to 2021 demonstrated that the new model reduced the average mean absolute error (MAE) evaluation metric by 18.9–30.06% for the prediction results of the 93 stations. The root mean square error (RMSE) metric decreased by 40.18–65.83%. For the prediction of maximum wind speeds exceeding level 6, the MAE was reduced by 40.41%, 25.93%, 19.96%, 21.39%, 12.39%, and 8.55% for the 6 forecast lead times, respectively. The RMSE evaluation metric also decreased by 30.92%, 18.67%, 12.29%, 12.21%, 7.92%, and 2.39% for the respective lead times. The improved model demonstrated consistent prediction performance and significantly enhanced accuracy.
基于地面实况观测数据、雷达组合反射率、风云四号A星波段数据 3 种实况观测资料,以 2018-2021 年广西 2850 个站点的小时累计雨量为预报对象,采用随机森林算法建立未来 1-3h的降水临近预报模型,并分别进行了单类型观测资料预报因子、3 种类型观测资料多种组合预报因子输入的预报试验.对各预报试验结果综合采用TS评分、命中率、虚警率和漏报率进行点对点的评估表明,地面资料在 1-3h的小雨和中雨量级、1-2h大雨量级的预报能力较好;雷达资料对未来 1h暴雨量级的预报较其他两种观测资料优势明显;卫星资料在小雨量级上有一定的预报能力,但其他量级各时效的预报均不理想.3 种观测资料在第 2-3h的暴雨预报能力都偏低.3 类预报资料因子组合预报结果的评估表明,各量级的多数预报时效均能比单类型资料预报取得更高的预报精度,其中大雨量级的预报提高了 10%以上,暴雨则提高了 25%以上.大雨和暴雨 1hTS评分的空间分布情况表明,组合因子高评分区域分布均最广,地面和雷达大雨量级的空间分布相当,雷达的暴雨TS评分在 0.2 以上空间分布范围较地面和卫星资料广,卫星的TS评分的空间分布显示其预报能力均最弱.
以1980-2020年广西台风期间桂林、梧州、龙州、南宁、玉林等5个气象观测站的地面日极大风速为研究对象,采用多元线性回归(MR)、支持向量机(SVM)、模糊神经网络(FNN)等三种较为常用的线性和非线性方法分别进行预报建模,对2011-2020年共10a独立样本的检验.结果 表明,在全样本风速预报的平均绝对误差上,FNN模型对桂林站、梧州站、龙州站、玉林站共4个站点预报的平均绝对误差最小,总体预报精度最好,MR预报模型则对南宁站有较好的预报能力,SVM模型预报效果总体偏差.对于6级以上大风的TS评分、命中率、空报率和预报偏差等4个评估指标的统计,FNN模型的预测精度最高且相对稳定,MR方案次之,SVM在三种方案中预报效果最差.FNN方法对广西台风期间地面日极大风速的预报有较好的参考作用.
为了更好地利用大量的卫星云图观测资料来提高台风暴雨的预报能力,解决并提高对台风强降水云系变化的预报精度,延长对未来云系变化的预报时效,构建基于合作对策Shapley-模糊神经网络的华南区域台风卫星云图非线性智能计算滚动集合预测模型,对增强卫星云图资料在台风暴雨天气预报中的实用性和及时性具有重要意义.依据2013—2016年华南区域台风影响过程的卫星云图,采用类似于数值预报模式的集合预报方法,通过对间隔6?h的卫星云图云顶亮温样本序列做经验正交函数分解,将提取出的时间系数作为云图预报建模的预报分量.考虑台风云系的发展变化主要受云团环境物理量场的影响,利用数值预报模式的物理量预报产品作为各预报分量的预报因子,并采用k-近邻互信息估计的分步式变量选择算法,通过两步过程实现相关变量的选择与弱相关变量的剔除,分别建立相应时间系数的Shapley-模糊神经网络集合预报模型,进一步将预报得到的各时间系数与空间向量合成,重构得到未来时刻的卫星云图预报图,实现了云图6—72?h的长时效客观滚动预测.试验结果表明,新方案所预测的云图与实况云图相关较高,重构云图的基本轮廓、纹理特征分布、清晰度以及云系强弱方面都比较接近原始云图.另外,研究进一步基于相同的云图预报因子,针对同样的建模和预报样本采用多元线性回归方案进行和新方案一致的云图预测.对比结果表明,这种非线性预报模型比线性方案能更好地预报未来较长时效台风云团的发展、移动的主要特征和变化趋势,其预测的云图与实际云图的主要特征更相似.云图预报时效达到了72?h,具有业务实用价值.
The recent emergence of satellite detection and imaging technologies has increased the demand for the application of satellite cloud images to current weather forecasting. However, approaches based on nonlinear prediction technology to forecast satellite images are lacking, and forecasting timelines are relatively short, e.g., 1–3 h in advance. In the present study, a nonlinear dimensionality reduction approach based on Laplacian eigenmaps (LEs) was combined with a random forest (RF) algorithm to construct an intelligent computing prediction model for rainstorm satellite images obtained from the first annual rainy season (April–June) in South China from 2010 to 2018. Results showed that the proposed forecasting model based on nonlinear intelligent calculation can accurately predict the key features and trends of the development and movement of strong precipitation clouds. The predicted satellite images described by the model were also consistent with the major features of the observed satellite images. This study then used a multiple linear regression (MLR) method based on the same prediction factors to establish a model for predicting satellite images for the same modeling and forecasting samples. Comparative results of the two prediction schemes showed that the LE + RF algorithm satellite image prediction scheme yields more samples exhibiting a high correlation with observed satellite images than the MLR method. Compared with that of the proposed scheme, the amount of samples of the MLR scheme in the low-correlation area was significantly larger. In general, the nonlinear intelligent computing scheme developed in this study is superior to the MLR method for predicting satellite cloud images. Thus, the LE + RF algorithm satellite image prediction scheme provides an objective and practical method for observed satellite cloud image predictions.
Rainstorm often causes inland flooding and mudslides that threaten lives and properties. In this study, rainstorm is used as a forecasting object, and an interpretation prediction model for rainstorm based on the European Center for medium-range weather forecasting (ECMWF) numerical prediction model is constructed through the generalized regression neural network method. Model inputs are forecasted through principal component analysis, and dual-factor feature extraction is performed on the primary predictors to obtain new irrelevant variables and optimize network structures. The experimental forecast results of the 24 h aging test using an independent sample of large-scale rainstorm in Guangxi, China from 2012 to 2016, the actual forecast results of selected rainstorm cases with great influence on Guangxi, and different influencing systems show that the new prediction scheme is sophisticated. Thus, the scheme has a certain universal applicability. The results of the comparative analysis between the new program and ECMWF show that the forecasting ability of the new method is more accurate than that of the direct numerical forecasting model. The threat score of the new forecast model for 5 years has a 58.4% increase relative to that of the ECMWF. The forecasting skills are positive and good and can thus improve the rainstorm forecasting ability of ECMWF and provide a better guidance for forecasters.
This study considers large-scale heavy rainfall as a forecast object based on the European central numerical forecast model product and uses a nonlinear fuzzy neural network (FNN) intelligent calculation method to establish a short-term forecast model of rainstorms. The information gain method is introduced to the predictor processing of the forecast model. Then the characteristics of many rainstorm predictors are calculated and screened on the basis of feature weight, information is condensed, some non-correlated forecast information variables are extracted, and the network structure of the forecast model is optimized. The modeled samples are determined and reconstructed by setting thresholds, and the modular forecast models of heavy rainfall and weak rainfall are established. The actual forecast results of the 24 h experimental prediction of the independent samples of large-scale rainstorms in Guangxi in 2012–2016 showed that the information gain-based modular FNN rainstorm forecasting model has higher prediction accuracy and a more stable forecasting effect. The various types of scores of 24 h of rainstorm (≧50 mm) at 89 weather stations in Guangxi from 2012 to 2016 are: threat score (TS) is 0.368, ETS: equal threat score (E) is 0.141, hit rate (POD) is 0.296, empty report rate (FAR) is 0.559, forecast bias (B) is 0.671, and HSS skill score (H) is 0.247. Further comparison and analysis of the European Centre for Medium-Range Weather Forecasts (ECMWF) numerical forecasting model forecast results indicated that the new model performed nonlinear intelligence calculated interpretation modeling on ECMWF numerical forecasting model products, and forecasting accuracy is improved to a certain extent compared with that of the original model. Forecasting techniques are positive and have good release effects, thereby improving the rain forecasting ability of ECMWF to a certain extent and providing a better reference value for business forecasters.
A nonlinear roiling prediction model for satellite image has been developed based on Shapley neural network using the ensemble prediction method similar to the numerical prediction model, due to lacking of the guidance of a nonlinear prediction theory for satellite image at present. Empirical Orthogonal Function(EOF) method is applied to the samples of infrared satellite image every 6 h in heavy rainfall processes, and time coefficients extracted are used as predictands. Since the changes of precipitation cloud system are governed by the physical quantity fields in cloud cluster, the physical quantifies prediction products from numerical prediction model are used as predictors, and Shapley Neural Network Ensemble Prediction models are established for the corresponding time coefficients based on the technique of the reduction of data dimensionality for data interpretation. By integrating the predicted time coefficient and space vector, the future satellite image is obtained. Results show that the nonlinear prediction model can better forecast the main features of the development of heavy rainfall cloud cluster in future 24h.
基于1961~2017年广西87个地面观测站逐日降水资料,利用NCEP/NCAP逐日再分析资料并综合运用诊断分析方法,从月际变化的角度分析广西前汛期大范围持续性暴雨的气候特征、大气环流特点以及水汽、动力等物理机制的差异.结果表明:(1)广西前汛期大范围持续性暴雨出现频数在4、5和6月份中呈逐月递增趋势.(2)不同月份发生大范围持续性暴雨的影响机制各异,500hPa表现为4月的两槽两脊并在低纬度地区有分裂出的短波槽影响广西;5月为两脊一槽形势;6月份的一槽一脊配合中低纬度的东亚槽.低层850hPa表现为异常的气流辐合,随着月份增加辐合不断加强.(3)4~6月的主要水汽来源和水汽含量各异.(4)4~6月广西上空不稳定能量增强,为广西暴雨的产生提供了有利的触发机制.