The effectiveness of the machine learning methods for real-world tasks depends on the proper structure of the modeling pipeline. The proposed approach is aimed to automate the design of composite machine learning pipelines, which is equivalent to computation workflows that consist of models and data operations. The approach combines key ideas of both automated machine learning and workflow management systems. It designs the pipelines with a customizable graph-based structure, analyzes the obtained results, and reproduces them. The evolutionary approach is used for the flexible identification of pipeline structure. The additional algorithms for sensitivity analysis, atomization, and hyperparameter tuning are implemented to improve the effectiveness of the approach. Also, the software implementation on this approach is presented as an open-source framework. The set of experiments is conducted for the different datasets and tasks (classification, regression, time series forecasting). The obtained results confirm the correctness and effectiveness of the proposed approach in the comparison with the state-of-the-art competitors and baseline solutions.
In modern data science, it is often not enough to obtain only a data-driven model with a good prediction quality. On the contrary, it is more interesting to understand the properties of the model, which parts could be replaced to obtain better results. Such questions are unified under machine learning interpretability questions, which could be considered one of the area’s raising topics. In the paper, we use multi-objective evolutionary optimization for composite data-driven model learning to obtain the algorithm’s desired properties. It means that whereas one of the apparent objectives is precision, the other could be chosen as the complexity of the model, robustness, and many others. The method application is shown on examples of multi-objective learning of composite models, differential equations, and closed-form algebraic expressions are unified and form approach for model-agnostic learning of the interpretable models.
In the paper, we discuss the applicability of automated machine learning for the effective multi-scale modeling of the industrial sensors time series. The proposed approach is based on the evolutionary generative design of the composite modeling pipelines. The iterative data decomposition algorithm is proposed in the paper to improve the quality of the sensor time series forecasting. To effectively use it in an automated way, the boosting-like mutation operators have been implemented for graphs-based genotypes. The proposed approach reduced the forecast error by 10% compared to the competitor library AutoTS. Also, the proposed modifications of the evolutionary algorithm resulted in better metrics in 78% of the cases where they were used.
In the paper, we propose an adaptive data-driven modelbased approach for filling the gaps in time series. The approach is based on the automated evolutionary identification of the optimal structure for a composite data-driven model. It allows adapting the model for the effective gap-filling in a specific dataset without the involvement of the data scientist. As a case study, both synthetic and real datasets from different fields (environmental, economic, etc) are used. The experiments confirm that the proposed approach allows achieving the higher quality of the gap restoration and improve the effectiveness of forecasting models.
The publication is considered methodology and preliminary results of computer processing of vector ice charts and simulation of ice navigation by empirical statistical model, which was designed in the AARI. The model is result of generalization of multi-year special ship ice observations, produced by the AARI scientists. The Institute has unique archive of ice charts, which were made by recognition of satellite images, their composition and vectorization. The archive covers the period since 1997. Average navigation velocity and total time expenditure for sailing along a route can be used as the objective indexes of the ice navigation conditions. In the article is presented detailed description of the technique for preparing data for numerical experiments with the empirical statistical model by applying the ice charts from the archive. The data preparing is processed in ArcGIS by means of specially designed computer programs. The numerical experiments simulate ice navigation of “Arctica” nuclear icebreaker. The comparison between the average navigation velocity and time expenditure for sailing along the route “Sabetta Port – Kara Gate Strait” at simulation of ice conditions in the first ten-day interval of May 1998 and the same temporal interval of 2019 indicates significant improve of the ice navigation conditions.
The paper presents a hybrid approach for short-term river flood forecasting. It is based on multi-modal data fusion from different sources (weather stations, water height sensors, remote sensing data). To improve the forecasting efficiency, the machine learning methods and the Snowmelt-Runoff physical model are combined in a composite modeling pipeline using automated machine learning techniques. The novelty of the study is based on the application of automated machine learning to identify the individual blocks of a composite pipeline without involving an expert. It makes it possible to adapt the approach to various river basins and different types of floods. Lena River basin was used as a case study since its modeling during spring high water is complicated by the high probability of ice-jam flooding events. Experimental comparison with the existing methods confirms that the proposed approach reduces the error at each analyzed level gauging station. The value of Nash–Sutcliffe model efficiency coefficient for the ten stations chosen for comparison is 0.80. The other approaches based on statistical and physical models could not surpass the threshold of 0.74. Validation for a high-water period also confirms that a composite pipeline designed using automated machine learning is much more efficient than stand-alone models.
In recent decades there has been a trend towards an increase in the number of dangerous hydrological events, especially floods. In order to protect citizens and solve economic problems, it is important to develop and actively introduce into operational practice methods of hydrological forecasting, as well as to build more modern and convenient interfaces of interaction between hydrometeorological services, municipal authorities and citizens. This work discusses a compact automated short-term hydrological forecasting system that uses open-source conceptual models HBV, SimHYD and GR4J as its core. The system is connected to data streams on the observed temperatures and precipitation in the watershed basin, as well as the predicted values of these parameters (in a current implementation, the WRF model with a forecast for 84 hours is used). Also, for operational calibration in daily mode, the system can assimilate (if available) data on observed water levels. Testing of the system is carried out on the example of Tikhvin city (the Tikhvinka river), which in recent years has been characterized by frequent flooding.
The article presents the results of the development of a model for calculating levels at one gauging station using the levels at another. To link the levels at two gauging stations, the data on levels, temperature and precipitation were used. The use of machine learning methods to solve the problem of predicting water levels made it possible to achieve an accuracy of about 6 cm. At the same time, traditional statistical models (linear regression, polynomial regression) have 14-16 cm error.
Satellite remote sensing has now become a unique tool for continuous and predictable monitoring of geosystems at various scales, observing the dynamics of different geophysical parameters of the environment. One of the essential problems with most satellite environmental monitoring methods is their sensitivity to atmospheric conditions, in particular cloud cover, which leads to the loss of a significant part of data, especially at high latitudes, potentially reducing the quality of observation time series until it is useless. In this paper, we present a toolbox for filling gaps in remote sensing time-series data based on machine learning algorithms and spatio-temporal statistics. The first implemented procedure allows us to fill gaps based on spatial relationships between pixels, obtained from historical time-series. Then, the second procedure is dedicated to filling the remaining gaps based on the temporal dynamics of each pixel value. The algorithm was tested and verified on Sentinel-3 SLSTR and Terra MODIS land surface temperature data and under different geographical and seasonal conditions. As a result of validation, it was found that in most cases the error did not exceed 1 °C. The algorithm was also verified for gaps restoration in Terra MODIS derived normalized difference vegetation index and land surface broadband albedo datasets. The software implementation is Python-based and distributed under conditions of GNU GPL 3 license via public repository.
Приводятся методика и результаты обработки векторных ледовых карт из архива ААНИИ за период 1998- 2018 гг. Получены ряды многолетней изменчивости суммарных протяженностей участков маршрута порт Сабетта - Берингов пролив в припае, в сплоченных льдах, в сплоченных льдах при наличии определенных возрастных категорий льдов и их частных концентраций, суммарной приведенной протяженности маршрута в старых и толстых однолетних льдах для десятидневных интервалов (декад) апреля и мая. Под термином «сплоченные льды» в статье понимаются дрейфующие льды общей сплоченностью 9, 9-10, 10 баллов за исключением случаев, когда они представлены исключительно начальными льдами толщиной до 10 см. Выполнена проверка рядов на наличие трендов методом интегральных кривых и проверка однородности рядов с помощью ранговых непараметрических критериев Уилкоксона-Манна-Уитни и Зигеля-Тьюки. Проанализировано более 4 тыс. значений протяженностей. Выявлено уменьшение суммарной протяженности участков маршрута в припае и в сплоченных льдах с наличием старых льдов, увеличилась протяженность пути в сплоченных льдах, в сплоченных льдах при наличии однолетних льдов средней толщины, в сплоченных льдах при наличии толстых однолетних льдов, в сплоченных льдах с частной концентрацией толстых однолетних льдов 5 и более баллов, в сплоченных льдах с суммой частных концентраций толстых однолетних льдов и однолетних льдов средней толщины 5 и более баллов. Уменьшение приведенной протяженности пути плавания в старых льдах частично компенсируется увеличением практически на эту же величину приведенной протяженности пути плавания в однолетних толстых льдах.
The article is considered methodology and results of computer processing of vector ice maps of the AARI archive for the period 1997-2018. The charts were produced by processing of Earth's remote sensing data. There were analyzed inter-annual variability of navigation conditions along two routes between the Sabetta Port (the Yamal Peninsula, Russia) and the Bering Strait during September that is annual period of the most light ice conditions. As the results of the processing there were obtained long-standing series of total lengths of the routes legs within free water, within close floating ice with concentration more than six tenths, and such integrated ice index as conditional lengths of the routes within compact floating ice with 10 tenths concentration. The series belong ten-days periods (decades) during September. The series were tested for the existence of trends by the method of integral curves, and then were examined for heterogeneity using Wilcoxon-Mann-Whitney and Siegel-Tukey rank non-parametric criteria. We received the following results: the total lengths of the routes legs within free water increased for the period of 1997-2018 years, ones within close floating ice and conditional lengths of the routes within compact floating ice decreased. There is some improvement of the ice conditions.
Results of testing of computer simulation model for assessment of probability of accidents with tankers due to pressure by drifting ice are presented. The testing was carried out for the navigation route «Sabetta Port – Kara Gate Strait – Murmansk Port» and for the first ten-days period of May, during the most difficult ice conditions of the navigation. The probabilities of the accidents were calculated. There was analyzed the model response to variations of its parameters values.
The paper discusses the methodology and results of electronic ice charts processing. The charts taken from AARI archive. The Barents, Kara, Laptev, East Siberian and Chukchi seas Ice maps reflect ice conditions for the period from 1997 to 2018 for the April-May inter-annual interval. The total stage lengths of «Sabetta – the Kara Gate –Murmansk» and «Sabetta – the Vilkitski Strait – the Bering Strait» standard routes were calculated at certain conditions of ice navigation. The route “Sabetta – the Bering Strait” was divided into sections within the Kara sea, Laptev Sea, East Siberian and Chukchi Seas for analysis. The purpose of the study is to obtain the values of the length of the routes in different categories of ice and to analyze changes trend of navigation in ice conditions for the period 1997-2018. The series were checked for the presence of trends using the integral curves method. The homogeneity of the series was checked using Wilcoxon - Mann-Whitney and Siegel - Tukey rank non-parametric criteria. Most of the series proved to be non-homogeneous. The following conclusions were made: there was some improvement of ice navigation conditions along the route Sabetta ‒ the Kara Gate – Murmansk due to the decrease of the route length in hard ice conditions. The ice navigation conditions along the Sabetta ‒ the Bering Strait route changed little, if at all, the navigation conditions along the route within the Kara Sea and the Laptev Sea have changed for the worse, and within the East Siberian Sea ice conditions scarcely changed. Some slight improvement of the navigation conditions was noted within the Chukchi Sea. In general, the decrease of the route Sabetta – the Bering Strait length in compact drift ice with total concentration equal to 9 tenths or more and in the presence of old ice is partially compensated by increase of the route length in compact ice in the presence of thick first-year ice. The decrease and the increase are relatively equal.
Start-up of plant for liquid gas production in Sabetta town (Russia, Yamal Peninsula) demands creation of the marine transport system for the gas shipment. The town is the significant cluster for the gas processing and export from South-Tambey field. The field stock is estimated in more than 1 trillion cubic meters of natural gas. Ice cover is a source of possible accidents with tankers. The accidents can lead to environment contamination. There is considered analysis of the intra- and inter-annual variability of ice conditions on the navigation route "Sabetta seaport – Kara Gate – Murmansk seaport" for the period of 1997-2018. The route is a path for export of liquid gas. It has processed electronic ice maps of the archives of the Arctic and Antarctic Research Institute (AARI). The maps of the archive were created by the vectorization of satellite images. The lengths of the route in close floating ice with presence of thick, medium, thin first-year ice, grey-white, grey and new ice, and with partial concentrations of thick and sum of thick and medium first-year ice equal to 5-10 tenths have been obtained for 24 decades from November to June. There has been carried out tests of homogeneity of the lengths inter-annual series by means of the method of integral curves and nonparametric Wilcoxon-Mann-Whitney and Siegel-Tukey criteria. Improvement of ice conditions along the route is revealed.