This paper investigates the problem of structuring knowledge in the field of inductive modeling. Inductive modeling algorithms are effective means of automatically constructing models based on experimental data. In this area, the theory of inductive modeling has been developed, new algorithms and software have been developed, and many applied problems in various areas of human activity have been successfully solved. The organization and structuring of knowledge on typical components of inductive modeling algorithms occur within the framework of the development of a metamodel for this area. A general representation of the metamodel is proposed, including sub-metamodels of individual components of the modeling process.
The paper investigates the problem of constructing a system of models for an object with many input and output variables for the case of modeling the dependence of mechanical properties of a metal casting on the chemical composition of the raw materials. A technique for constructing system models for multidimensional objects from given experimental data is described The initial data set contains numerical values of five input variables (carbon, silicon, manganese, chromium, and phosphorus) and four output variables (tensile strength, relative elongation, impact strength, and hardness) characterizing the iron casting. Based on these data, the system of models is obtained describing the dependence of the output vector of casting properties on the input vector of its chemical composition and allowing to model these properties instead of conducting expensive experiments. To construct the system model, a technology of building models of multidimensional objects is applied based on the combinatorial GMDH algorithm allowing automatically derive linear or nonlinear models from a set of experimental data. As a result of modeling, a nonlinear system model of optimal complexity is obtained that allows explaining the dependence of the casting properties on chemical composition and the interdependence of these properties between them. This system model can be used for choosing the relevant chemical composition of raw materials for the specified physical and mechanical properties of the iron casting.
The paper solves the problem of discovering the dependence of mechanical properties of a metal casting on the chemical composition of source materials by constructing a system of models for an object with many input and output variables. For this purpose, a system of models for output variables is constructed using the system technology based on the combinatorial GMDH algorithm. Such a system of models will allow simultaneously choosing the chemical composition of the raw material ensuring the specified mechanical properties.
The paper investigates the problem of constructing a system of models for an object with many input and output variables using the example of modeling the dependence of the mechanical properties of a metal casting on chemical composition of raw materials. A technique for constructing system models for multidimensional objects based on experimental data is described. As a result of modeling, a system of models was obtained that allows one to simultaneously select the chemical composition for given mechanical properties.
The problem of constructing a baseline for the multicomponent signal of the intensity of inverse chronopotentiometry, the determination of which allows to estimate the concentration of various chemical elements dissolved in water quite accurately, is investigated. To solve this problem, an approach to construction of the approximation function of the lower envelope line of the differential signal in different classes of basic functions with the use of GMDH is proposed. The approach was used for constructing the best model of the differential signal baseline on the real example of measuring the Zn concentration under the presence of ions Cd, Pb, Cu. The built model of optimal complexity is the sum of arguments with the direct and inverse degrees which is necessary for clearing the intensity signal from background to obtain the intensity spectrum of the measured chemical elements.
Computer technology of analysis, integral assessment and forecasting of the state of economic security of the state has developed for information support of making effective managerial decisions. The developed computer technology, designed to solve tasks in the field of economic security based on the integration of common software products Microsoft Excel and StatSoft Statistica, ensures its effective use by individuals and bodies making decisions in that field. This software product we developed to ensure automation of the process of economic security monitoring and assessment of the impact of various factors on its level.
The developed technology based on the GMDH was applied to solve the problem of choosing the best model that describes the dependence of the cooling temperature. The use of this technology increases the effectiveness of decision support for the caster in the process of manufacturing castings. The selection of the optimal cooling mode plays a significant role in obtaining a high-quality end product. To build models, the technology incorporates a hybrid iterative-genetic GMDH algorithm, the peculiarity of which is that it is a neural network algorithm with active neurons. The article compares two algorithms for constructing a cooling temperature model depending on the modes of the cooling installation: the LASSO algorithm and the hybrid iterative-genetic algorithm of the GMDH.
Inductive modeling methods have repeatedly demonstrated their effectiveness in various problems of building models based on experimental data. Many different GMDH algorithms of the sorting and iterative types are known which have already proved their efficiency, but they are constantly being improved. To date, some new effective GMDH algorithms have been developed, for example, combinatorial-genetic, correlation-rating and hybrid iterative-genetic algorithms. The paper compares the results of applying various GMDH algorithms with the classical and Lasso regression methods to the problem of modeling the cooling process of a metal casting.
Introduction.Data volumes are permanently increasing and some new approaches are needed for storage and processing them considering the development and improvement of modern computers.This puts forward new requirements to automatic data processing tools and intelligent systems for analyzing information with taking into account its semantics.The advantage of iterative group method of data handling (GMDH) algorithms is that they are able to work with a large number of arguments.The generalized iterative GMDH algorithm includes various former modifications of these algorithms.For example, algorithms of multilayer and relaxation types as well as varieties of iterative-combinatorial (hybrid) algorithms are diverse particular cases of the generalized one.Metamodeling is the construction of generalized models of a certain group of objects (software tools, mathematical models, information systems).An ontological metamodel of the iterative GMDH algorithms was built using the Protege tools in order to structure knowledge in this subject area. The purpose of the paper is to analyze the developed iterative GMDH algorithms and propose an approach to structuring knowledge оn iterative GMDH algorithms by building an ontological metamodel of this subject area.Results.A retrospective analysis of the developed iterative GMDH algorithms іs carried out in the paper, their advantages and disadvantages are indicated.It is shown that the generalized iterative algorithm, whose special cases are both known and new varieties of multilayer, relaxation and iterative-combinatorial GMDH algorithms, makes it possible to compare the effectiveness of various algorithms and solve real modeling problems.Based on the results of this study, an ontological metamodel of iterative GMDH algorithms has been developed.
Usually genetic algorithm (GA) uses crossover and mutation operator for solving optimization problems. But recent studies show that sometimes GA can be just as efficient using only one of these operators. In this paper, two-parametric mutation operator is proposed to generate models in combinatorial-genetic algorithm (COMBI-GA) when solving inductive modelling tasks. The algorithm efficiency is compared when using a pair of a crossover and a(0→1)-bit mutation or only two-parametrical (0↔1)-bit mutation when solving a real-world problem.
The paper deals with the problem of restoring missing experimental data in modeling tasks. Real data samples may have missing values for some variables or a study period. All this leads to the risk of building an inaccurate model and, as a result, the experiment failed. In the paper, this problem is considered on data of a real task. A non-linear interpolation model was built for the dependence of the concentration of chlorophyll in algae on the concentration of the pollutant dependent on time. The found model made it possible to restore missing data for the days when measurements were absent, as well as to find out on which day and at what concentration of the pollutant the algae would die.
Construction of an ontological metamodel of iterative algorithms is proposed to structure knowledge about these algorithms and their implementation. The metamodel will automate the design and use of specialized software for solving specific applied modeling problems. The advantage of iterative GMDH algorithms over combinatorial ones is that they allow the big datasets processing. The known generalized iterative algorithm, allows you to create typical architectures of previously developed modifications of these algorithms when setting up various modes of operation of this algorithm. The authors have developed an ontological metamodel of iterative GMDH algorithms using the Protégé environment.
Introduction.Deep neural networks are effective tools for solving actual tasks such as data mining, modeling, forecasting, pattern recognition, clustering, classification etc.They differ with respect to the architecture design, learning methods and so on.Most simple and widely used are deep feed-forward supervised NNs. The purpose of the paper is to compare briefly main features of the deep feed-forward deterministic supervised networks with the Multilayered Iterative Algorithm of GMDH (MIA GMDH) and to formulate main ideas of constructing a new class of hybrid deep networks based on the MIA neural network.Methods.Most usable deep feed-forward supervised neural networks have been studied: multilayered perceptron, convolutional NN and some its modifications, polynomial neural networks, genetic polynomial neural network etc.Results.There was carried out a comparative analysis of main features of the MIA GMDH neural network with the characteristics of other deep deterministic supervised neural networks.The most promising approaches are identified to improve the performance of this network, particularly by hybridization with methods of computational intelligence.The main idea of building a new class of hybrid deep networks based on MIA GMDH is formulated.
The task of structure designe of the software complex of tools for inductive modeling on the GMDH basis is considered. A novel is the use of the knowledge base in the form of an ontology of the subject area of inductive modeling. The application of the ontological approach to the construction of the knowledge base makes it possible to reuse information and effectively process this information when modeling complex systems of different nature using statistical data, to form queries and obtain logical conclusions. Fragments of the ontology of inductive modeling based on GMDH as an example of creating a formal description of the subject area are given. The Protégé_4.3 ontoeditor was used to construct ontological models of the inductive modeling domain.
A comparative analysis of different methods for filtering the measured noisy cooling curves of gray cast iron melts is carried out. Three different methods of noise filtering in the data are compared: simple moving average method as a baseline, method of adaptive smoothing as well as the proposed approximative polynomial filtration based on selection of the degree of the polynomial using GMDH. All this methods give a possibility to filter out noise in the data but the moving average and adaptive smoothing methods give only numerical filtration results. In contrast, the result obtained using the method of approximative filtration is a polynomial function which proved to be the most accurate and to keep all the critical points of the initial cooling curve.
The paper investigates the asymptotic convergence of some typical criteria for model selection from a given data sample. A range of known criteria are generalized into a special class joining two different groups based on both explicit and implicit implementing the trade-off between model accuracy and complexity. Criteria of the first group contain various explicit penalty terms for the model complexity whereas those from the second group are based on the cross-validation idea like the sample division into two parts which is typical for the GMDH criteria. A new approach to the analysis of asymptotic properties is introduced based on the assumption of strong regularity of regressors. This means that values of those inputs are not vanishing functions at infinity but are the square summable ones. Definitions of asymptotic characteristics of the generalized criterion are given and analyzed, namely convergence in probability as well as consistency. The investigation demonstrates that the strong regularity of regressors gives the sufficient condition for the convergence in probability of the generalized criterion to some finite values. Moreover, the minimal value of this generalized criterion at infinity corresponds to the true model manifesting that this criterion is consistent together with all the criteria of this class.
The article presents the modelling results of the Ukraine Covid-19 epidemy process using the combinatorial-genetic method based on official statistical data. A comparison with some other known methods for model construction and the process prediction is also given. This research is important for defining tendency of coronavirus evolvement in time and predicting its future activity in order to take some protective measures.
This article presents the modeling results of the Ukraine Covid-19 pandemic process using official statistical data on the confirmed cases. Main goal is discovering dynamic regularities of the process given as daily data of the time series. That is why we use four different methods to build predictive difference models of the autoregression type: ordinary autoregression; autoregression of optimal structure obtained using the combinatorial-genetic GMDH algorithm COMBI-GA; another variant of the optimal structure built by the well-known method Lasso; and we compare prediction results of these methods with independent predictions published by the World Data Center. In our study, the baseline prediction is the one produced by the ordinary autoregression which includes all lags from 1 to a given their number. Unlike it, algorithms COMBI-GA as well as Lasso construct autoregressions with in some respect optimal compositions of lags. The WDC independent predictions are made using the Backpropagation ANN as nonlinear transformation of lag variables. This comparative study we carried out in two stages: first, for the period of strong quarantine in Ukraine and second, for the period of stepwise quarantine weakening. For both stages, the optimal models built by the COMBI-GA are most interpretable and demonstrate better predictive accuracy on validation datasets. This research is useful for defining tendency of coronavirus evolvement in time and predicting its future activity in order to take some protective measures.
The main stages and components of a technology to solve problems of integrated (composite) evaluating and forecasting the performance of a complex economical system are described to support making efficient managerial decisions based on an aggregate and therefore observable information. We introduce the three novelties in the task of construction the synthetic integral index: some kind of fuzzy interpretation of values of a primary economic indicator; based on it a nonlinear method for normalizing indicators; approach to forecasting the index. An example of using the technology to analysis and forecast of the investment activity in Ukraine is presented.