The development and use of scientific applications have become an integral part of conducting large-scale experiments in various fields of research that require high-performance computing and big data processing. In the context of developing such applications, non-trivial problems arise in the concerted description and further use of schemes, software, and computational resources to solve subject domain problems of a specific application. Research productivity has become highly dependent on the degree of automation in the preparation and execution of experiments in a computing environment whose resources may be distributed and heterogeneous. Many approaches to the experiment automation are based on workflows as a structure for formalizing and specifying data processing and high-performance computing using distributed applications. Within such approaches, developers and end-users work with workflow management systems for the collaborative development and use of distributed scientific applications. Nowadays, service-oriented applications are coming to the fore. However, there is a wide range spectrum of problems related to the support of modular scientific applications, the standardization of their components and interfaces, the use of heterogeneous information and computing resources, and organization of interdisciplinary research within service-oriented architecture. Known workflow management systems do not fully address the above problems. In this regards, we consider relevant aspects of organizing service-oriented computing in a heterogeneous distributed computing environment. We propose a new framework for creating service-oriented and workflow-based scientific applications. The paper shows that the proposed framework significantly extends and complements the capabilities of systems for such purposes. We also demonstrate the reduction in labour costs associated with the preparation and execution of experiments.
Рассматривается проблема повышения эффективности управления вычислениями при выполнении научных приложений в гетерогенной распределенной вычислительной среде, схемы решения задач в которых реализуются научными рабочими процессами. Одним из ключевых факторов, влияющих на время выполнения рабочих процессов, является время, затрачиваемое на передачу данных между их модулями. В известных системах управления рабочими процессами передача данных между модулями зачастую осуществляется через центральный узел, что увеличивает время выполнения этих процессов. В работе предложены средства управления вычислениями с учетом структуры рабочего процесса и передачи данных напрямую между узлами. Разработан эвристический алгоритм назначения узлов для запуска модулей рабочего процесса на основе оценки времени их выполнения и размера данных, передаваемых между ними. Результаты экспериментов показывают существенное сокращение времени выполнения ряда известных научных рабочих процессов при использовании данного алгоритма относительно применения алгоритмов подобного назначения в системе HTCondor.
Предложен новый подход к созданию сервис-ориентированной вычислительной среды для моделирования и оценки функционирования и развития природно-технических систем с использованием высокопроизводительных вычислительных систем. Представлена схема прогнозирования развития и оценки состояния локального энергетического комплекса. Апробация предложенного подхода выполнена на примере решения задачи комплексной оценки критичности элементов транспорта локальной системы газоснабжения.
Nowadays, multi-energy systems play an important role in satisfying the ever-increasing demand for different energy resources. At the same time, the sustainable development of such systems is usually based on the structural and parametric optimization (synthesis) of their infrastructures. There is a large spectrum of specialized optimization tools for the study of single energy systems. At the same time, the problem of modeling the interaction between single energy systems remains challenging. Therefore, it is imperative to develop an efficient experimental environment to effectively implement the synthesis of optimal configurations of multi-energy systems. Microgrids are a special case of multi-energy systems. They provide a higher level of energy supply compared to the main grids and enhance their reliability and resilience. In this context, we propose a framework and subject-oriented environment for the synthesis of optimal microgrid configurations in a reasonable time considering the available computational resources. The basis of the environment is a service-oriented application. The modeling and optimization of the studied systems is performed by means of scientific workflows. Our results complement and develop known approaches to automate the modeling of multi-energy systems using their typical models and specially selected optimization algorithms corresponding to these models. We have successfully tested our approach for the synthesis of optimal microgrid configurations on the case study of a specific microgrid providing heat and electricity to a small settlement.
The nature of multi-energy systems requires the integration of interdisciplinary methodologies and complementary tools in a coherent and systematic manner. Therefore, it is imperative to develop an efficient experimental environment to effectively implement the interaction between different tools for structural and parametric optimization (synthesis) of multi-energy systems and analysis of their operation. Due to the inherent complexity of the energy system optimization model, solving the synthesis problem often leads to significant challenges in terms of required computational resources and runtime. Known synthesis approaches do not fully address these challenges. In this context, we propose a subject-oriented environment for the multi-energy system synthesis in a reasonable time considering the available computational resources. The basis of the environment is a service-oriented application. The modeling and optimization of the studied systems is performed by means of scientific workflows. The execution of the specialized tools for modeling and optimization is done through web services. The pre-selection of these tools, depending on the level of model detail of the systems under study, and the data validation are performed on the specially developed testbeds. Such testbeds are constructed as system workflows. This is one of the distinguishing features of our approach. The application of the proposed environment is demonstrated through a case study on the synthesis of a microgrid (a special case of a multi-energy system). As a result of the study, an optimal microgrid configuration was proposed. It is based on the cogeneration of heat and electricity generation using natural gas.
We propose the agent-based approach to intellectualize data processing and analysis in modeling the operations of interconnected microgrids. Microgrids are modern energy systems with a large share of environmentally friendly and resource-saving equipment. The modeling of different aspects for the organization and operation of microgrids is relevant for their structural and parametric optimization. The study of these aspects in the interaction of microgrids deserves special attention. In this case, it becomes possible to take into account the synergistic effect in the distribution and consumption of electric power. We implement the modeling process using a multi-agent system. In this regard, we describe the structure of this system, methods of its construction, and goals and processes of agent operations. In addition, we provide a model of agent behavior developed based on the basis of a finite-state control machine. The advantages of the proposed approach are demonstrated on a model example of balancing generation and consumption of electric power.
Стремительное развитие параллельных и распределенных вычислительных систем, телекоммуникационных технологий и облачных платформ обеспечило возможность разработки и применения научных приложений с целью подготовки и проведения крупномасштабных экспериментов с большими массивами данных. Зачастую создаваемые приложения предполагают сложную схему решения задач, базирующуюся на интегрированном выполнении процессов передачи, обработки и анализа данных, ресурсоемких вычислений и принятия решений. При этом математическое и программное обеспечение приложений может реализовываться различными группами специалистов из разных организаций и ориентироваться на разнородные вычислительные ресурсы. Все это обусловливает необходимость использования развитых средств проектирования, реализации, развертывания и выполнения научных рабочих процессов в рамках единой распределенной вычислительной среды, интегрирующей в конечном счете алгоритмические знания, программно-аппаратное обеспечение, данные и разнообразные сервисы. В настоящее время в качестве таких средств, как правило, выступают системы управления рабочими процессами. В этой связи статья посвящена обсуждению текущего состояния известных систем управления рабочими процессами, а также рассмотрению проблем, связанных с научными рабочими процессами для различных вычислительных сред. Актуализируются проблемы разработки и применения подобных систем, которые в настоящее время не решены в полной мере. В частности, отмечается необходимость учета предметной специфики и обеспечения масштабирования вычислений, востребованность сервис-ориентированных приложений, эффективность эксплуатации гетерогенных сред, интегрирующих высокопроизводительные ресурсы пользователей, кластерные ресурсы центров коллективного пользования, Grid-систем и облачных платформ.
Рассматривается подход к разработке и применению испытательного стенда сервис-ориентированных приложений в гетерогенной распределенной вычислительной среде. Целью испытательного стенда является обеспечение разработчиков средствами проведения экспериментов по оценке качества данных, работы алгоритмов, анализа результатов вычислений и других аспектов реализации сервис-ориентированных приложений. Сервис-ориентированное приложение создается на основе рабочих процессов, представляемых в виде композиции веб-сервисов. Веб-сервисы реализуют выполнение прикладного и системного программного обеспечения. Прикладное программное обеспечение поддерживает выполнение операций предметной области приложения. Суть и новизна подхода заключается в использовании в рамках испытательного стенда системного программного обеспечения, предоставляющего дополнительные функции обработки и анализа данных, получения информации с контрольно-измерительного оборудования и многокритериального выбора оптимальных результатов вычислений. Рабочие процессы могут быть представлены в стандартной форме на декларативном языке Business Process Execution Language для использования во внешних системах управления рабочими процессами.
Implementing high-performance computing (HPC) to solve problems in energy infrastructure resilience research in a heterogeneous environment based on an in-memory data grid (IMDG) presents a challenge to workflow management systems. Large-scale energy infrastructure research needs multi-variant planning and tools to allocate and dispatch distributed computing resources that pool together to let applications share data, taking into account the subject domain specificity, resource characteristics, and quotas for resource use. To that end, we propose an approach to implement HPC-based resilience analysis using our Orlando Tools (OT) framework. To dynamically scale computing resources, we provide their integration with the relevant software, identifying key application parameters that can have a significant impact on the amount of data processed and the amount of resources required. We automate the startup of the IMDG cluster to execute workflows. To demonstrate the advantage of our solution, we apply it to evaluate the resilience of the existing energy infrastructure model. Compared to similar approaches, our solution allows us to investigate large infrastructures by modeling multiple simultaneous failures of different types of elements down to the number of network elements. In terms of task and resource utilization efficiency, we achieve almost linear speedup as the number of nodes of each resource increases.
Статья посвящена разработке инструментальных средств построения гетерогенной вычислительной среды для создания и эксплуатации распределенных пакетов прикладных программ. Реализация научных рабочих процессов (схем решения задач) в рамках пакета основывается на контейнеризации прикладного и системного программного обеспечения. Суть и новизна представленного подхода заключаются в совместном, согласованном использовании стандартизированного языка Business Process Execution Language, веб-сервисов обработки данных и контейнеризации научных рабочих процессов. Такая интеграция позволяет: описать на вышеупомянутом языке схемы решения задач, объединяющие сервисы для доступа к базам данных и приложениям предметных специалистов из разных организаций; обеспечить за счет контейнеризации масштабирование как процесса выполнения вычислительных заданий, так и процесса функционирования систем управления этими заданиями в среде; осуществить взаимодействие научных рабочих процессов с геоинформационными системами с помощью веб-сервисов обработки данных. Результаты исследования успешно апробированы для пакета анализа уязвимости энергетических комплексов и отдельных систем энергетики на территории Байкальской природной территории. The paper addresses the design of tools for constructing a heterogeneous computing environment for the development and use of distributed applied software packages. The implementation and execution of scientific workflows (problem-solving schemes) within the package is based on containeri- zation of applied and system software. The essence and novelty of the represented approach lies in the integrated use of the standardized Business Process Execution Language, web processing services and containerization in implementing and executing scientific workflows. This integrated use allows describing the aforementioned language problem-solving schemes that combine services for accessing databases and applied software, which are developed by subject matter experts from different organizations. To ensure scaling of both the process for executing computational jobs and the functioning of systems for managing these jobs in a heterogeneous computing environment through containerization the proposed approach was applied. Finally, the described method was used to integrate scientific workflows with geographic information systems using web processing services. The results of the study were successfully applied for the development and use of a package for analyzing the vulnerability of energy complexes and separated energy systems in the Baikal Natural Territory.
The solution of environmental monitoring problems reasonably requires the collection, digitization, storage, and analysis of a large volume of spatiotemporal data. Their processing frequently necessitates the application of distributed calculations. The known geoinformation systems do not usually contain the means, which could support the calculations of this kind to a full extent. An approach to the solution of this problem is proposed on the basis of a geoinformation system integrated with the means of the development and use of distributed scientific applications presented by web data processing services in a heterogeneous computational environment. An advantage of this approach consists in a decrease in the labour efforts in the preparation and implementation of experiments with geodata.
Статья посвящена проблеме интеллектуализации обработки и анализа данных в исследовании процессов функционирования микросетей как экологически чистых и ресурсосберегающих систем энергетики. Процесс моделирования взаимодействия микросетей реализуется мультиагентной системой. Рассмотрены структура мультиагентной системы, средства ее построения, назначение и процессы функционирования агентов. Разработана модель поведения агента, базирующаяся на использовании конечного управляющего автомата. Мультиагентная система ориентирована на поддержку исследования живучести автономных энергетических систем инфраструктурных объектов Байкальской природной территории. A microgrid is a network having a diverse spectrum of energy generators and storage devices. Generally, it has a relatively small local loads. Energy management in such networks has a number of advantages, such as reducing power losses and simplifying the management process. To this end, the need to study various aspects of organizing and operating microgrids from the point of view for managing their functioning is significantly increases. In this regards, the paper addresses an intellectualization of data processing and analyzing within the study of micro-grids as environmentally friendly and resource-saving energy systems. We consider microgrid interactions under external disturbances. The process of modelling the interaction of microgrids is realized by a multi-agent system. The multi-agent system structure, tools of its construction, and purpose and functioning of agents are considered. Moreover, we developed and presented a model of agent behavior that is based on the use of a finite control machine. The multi-agent system is implemented using the JADE framework. To reduce labor costs for creating agents, we have developed a special add-on to JADE for determining the agent behavior. The considered multi-agent system is focused on supporting the study of the resilience of autonomous energy systems for infrastructure objects located on the Baikal natural territory.
Nowadays, the development and use of workflow-based applications (distributed applied software packages) are some of the key challenges in terms of preparing and carrying out large-scale scientific experiments in distributed environments with heterogeneous computing resources. The environment resources can be represented by clusters of personal computers, supercomputers, and private or public cloud platforms and differ in their computational characteristics. Moreover, the composition and characteristics of resources change in dynamics. Therefore, computations planning and resource allocation in the considered environments are important problems. In this regard, we propose new algorithms for computation planning taking into account redundancy and uncertainty in such distributed applied software packages. Compared to other algorithms of a similar purpose, the proposed algorithms use evaluations of workflow execution makespan obtained in the process of continuous integration, delivery, and deployment of applied software. The proposed algorithms provide the construction of redundant problem-solving schemes that allow us to adapt them to the dynamic characteristics of computational resources and improve distributed computing reliability. The algorithms are based on a theory of conceptual modeling computational processes. We demonstrate the process of constructing problem-solving schemes on model examples. In addition, we show the utility in using redundancy for increasing the distributed computing reliability In comparison with some traditional meta-schedulers.
Nowadays, data analysis is an integral part of large-scale scientific and applied experiments. Opportunities of modern computing environments allow us to move from operating with traditional storage systems within solving data-intensive problems to the in-memory data grid technologies. Such technologies improve the performance and scalability of data processing compared to external databases because of a faster random access memory and other hardware advancements. The considered data grid enables applications to cache data in the memory. Based on our practical experience, we discuss the advantages of applying the in-memory data grid technology to analyze the energy system vulnerability. The complexity of this problem is determined by considering possible combinations of simultaneous failures of energy system elements. We use open source-based Apache Ignite to support high-performance computing and data distribution. The study aims to evaluate the impact of the data grid scaling on the problem-solving quality criteria. We used the resources of the public access Irkutsk Supercomputer Center to carry out experiments.
Nowadays, developing and applying advanced digital technologies for monitoring protected natural territories are critical problems. Collecting, digitalizing, storing, and analyzing spatiotemporal data on various aspects of the life cycle of such territories play a significant role in monitoring. Often, data processing requires the utilization of high-performance computing. To this end, the paper addresses a new approach to automation of implementing resource-intensive computational operations of web processing services in a heterogeneous distributed computing environment. To implement such an operation, we develop a workflow-based scientific application executed under the control of a multi-agent system. Agents represent heterogeneous resources of the environment and distribute the computational load among themselves. Software development is realized in the Orlando Tools framework, which we apply to creating and operating problem-oriented applications. The advantages of the proposed approach are in integrating geographic information services and high-performance computing tools, as well as in increasing computation speedup, balancing computational load, and improving the efficiency of resource use in the heterogeneous distributed computing environment. These advantages are shown in analyzing multidimensional time series.
Nowadays, the development of scalable scientific and applied applications is one of the key challenges for modern distributed environments that include heterogeneous computing resources. Clusters of personal computers, high-performance systems of public access supercomputer centers, or cloud platforms often represent these resources. Usually, cluster nodes of different types differ significantly in their computational characteristics (the number of nodes and cores, RAM and disk memory, interconnect, system software, administrative policies, job scheduling principles, and many other parameters). Therefore, planning computations in distributed applied software packages is an important problem. To this end, we propose new models and algorithms for computation planning in software packages. Their distinctive feature is the analysis of the execution time of package modules in the environment nodes. We obtain such time evaluations in the process of their continuous integration, delivery, and deployment. The proposed models and algorithms allow us to create redundant problem-solving schemes. In practice, the problem-solving scheme redundancy enables us to adapt it to the changing characteristics of environment resources and improve the computation reliability. Construction of problem-solving schemes is shown in model examples.
Nowadays, simulation modeling is a relevant and practically significant means in the field for research of infrastructure object functioning. It forms the basis for studying the most important components of such objects represented by their digital twins. Applying meteorological data, in this context, becomes an important issue. In the paper, we propose a new microservice-based approach for organizing simulation modeling in heterogeneous distributed computing environments. Within the proposed approach, all operations related to data preparing, executing models, and analyzing the obtained results are implemented as microservices. The main advantages of the proposed approach are the parameter sweep computing within simulation modeling and possibility of integrating resources of public access supercomputer centers with cloud and fog platforms. Moreover, we provide automated microservice web forms using special model specifications. We develop and apply the service-oriented tools for studying environmentally friendly equipment of the objects at the Baikal natural territory. Among such objects are recreation tourist centers, children's camps, museums, exhibition centers, etc. As a result, we have evaluated the costs for the possible use of heat pumps in different operational and meteorological conditions for the typical object. The provided comparative analysis has confirmed the aforementioned advantages of the proposed approach.
A new approach to creating a domain-specific heterogeneous distributed computing environment is considered. It is used to support decision-making on urgent problems of increasing the survivability of energy systems. The approach is based on the use of high-performance computing, multiagent scheduling of computations and resource assignment, tools for processing semistructured information, and visualization of subject data using electronic maps. Decision-making alternatives are evaluated using combinatorial modeling and multicriteria optimization. Orlando Tools is used as the base for the environment's integrated software. It implements flexible modular construction of scalable scientific applications (distributed application packages). The advantages of using the environment are demonstrated by the example of solving practical problems.