Accurate error handling in user data is a crucial component for the integrity and reliability of information systems. Errors can arise from a myriad of sources, such as typos introduced by the user, misinterpretations of the input fields intended content, or general user confusion about the specifications of the required input, among others. These errors, if not rectified, can lead to significant data inaccuracies and, as a consequence, flawed decision-making. An advanced perspective on improving error handling involves the intricate task of pinpointing hidden semantic relationships within the data. These subtle, implicit connections may not be readily salient but hold immense potential in influencing the data’s integrity. Detecting these relations involves careful analysis of the context in which the data is used or employing cutting-edge semantic algorithms designed to uncover relationships that escape the naked eye. This paper delves into the fundamental principles and outlines a strategic roadmap for the implementation of error handling in user data within the Octoshell supercomputer center management system. It aims to dissect the crucial steps necessary for the incorporation of robust error detection and correction mechanisms that are integral to maintaining data integrity.
Increasing the utility output of supercomputer centers is the subject of many studies due to the high cost of operating such systems. Regardless of whether access to supercomputer systems is provided commercially or free of charge, the issue of their efficient usage is of great importance for system holders. Optimizing the queues of jobs and working with custom applications are just some of the opportunities to increase the efficiency of such large computers. This work examines the possibilities of reducing the time until users receive results by reducing the waiting time in the queue and informing users and administrators about the changing behavior of the series of job runs. The proposed approach is being tested at the supercomputer complex of Moscow State University, and the Lomonosov-2 supercomputer.
Supercomputer co-design provides for coordination of the specifics of solved problems, methods, algorithms and their implementations with the architecture of computing systems. In this case, supercomputer co-design technologies usually provide for the solution of the following tasks: search for an optimal method or algorithm for solving a certain problem on a particular supercomputer, construction of an optimal implementation of the selected method and algorithm on the target computer system, as well as the selection or construction of the most appropriate hardware and software platform for solving the problem using the selected method or algorithm on the basis of this or that software implementation. All three variants of the scientific problem statement are meaningful, require different approaches for solution, but have in common—the necessity of joint analysis of properties of problems, methods, algorithms, software implementations and computer architectures for effective solution.
The paper discusses the development of the Algo500 scalable digital platform project. One of its main components is the PerfData data repository, which includes all the parameters necessary to perform a supercomputer experiment, as well as the results obtained on it. All data entered into PerfData must be clearly defined, well-structured and unambiguously understood by various researchers. The purpose of this paper is to describe these data for their subsequent use in the Algo500 project.
Effective output from data centers are determined by many complementary factors. Often, attention is paid to only a few, at first glance, the most significant of them. For example, this is the efficiency of the scheduler, the efficiency of resource utilization by user tasks. At the same time, a more general view of the problem is often missed: the level at which the interconnection of work processes in the HPC center is determined, the organization of effective work as a whole. missions at this stage can negate any subtle optimizations at a low level. This paper provides a scheme for describing workflows in the supercomputer center and analyzes the experience of large HPC facilities in identifying the bottlenecks in this chain. A software implementation option that gives the possibility of optimizing the organization of work at all stages is also proposed in the form of a support system for the functioning of the HPC site. Эффективная отдача от вычислительных центров определяется множеством взаимодополняющих факторов. Зачастую внимание уделяется только нескольким, на первый взгляд наиболее значимым из них. Например, это эффективность работы планировщика, эффективность использования ресурсов пользовательскими задачами. При этом часто упускается более общий взгляд на проблему: уровень, на котором определяется взаимосвязь рабочих процессов в СКЦ, организация эффективной работы в целом. Упущения на этом этапе способны обесценить любые тонкие оптимизации на низком уровне. В данной работе приводится схема описания рабочих процессов в СКЦ, анализируется опыт крупных СКЦ по выделению наиболее узких мест в этой цепочке, предлагается вариант программной реализации, дающей возможность оптимизации организации работ на всех этапах, в виде системы поддержки функционирования СКЦ.
Users, while developing and executing their applications, can choose from several supercomputer systems to get computation results. This choice is influenced by many factors, and the wait time before application execution is among them, because users can not immediately get the required resources which are allocated for computing the results of other users’ computational jobs. Special software, called schedulers, is used to manage and share resources. Although schedulers use fixed algorithms for this task, the queue wait time is often difficult to estimate, because of the lack of information about jobs and future events. In this research, different methods of job queue wait time estimation are compared, some algorithms are modified and evaluated, using execution logs of the ‘‘Lomonosov-2’’ supercomputer.
Supercomputer is an exceptionally valuable computational resource and it must be used as efficiently as possible. However, in practice, the efficiency of its usage leaves much to be desired. There are various reasons for this. One of the main ones is the low performance of user applications, but users themselves are often not aware of the presence of performance issues in their programs. Therefore, it is necessary for administrators of a supercomputer to be able to constantly monitor the performance and behavior of all running jobs. However, the problem is that the commonly used metrics for assessing the quality of resource consumption (such as CPU or GPU load, the amount of bytes transferred over the MPI network, etc.) are often far from being convenient and accurate. This paper describes the implementation and evaluation of the previously proposed assessment system, which, in our opinion, makes it possible to significantly ease the task of properly evaluating the quality of the supercomputer resource usage. We also touch upon another topic related to the assessment of the quality of using HPC resources — organization of HPC resource provisioning.
Supercomputer technologies are in demand for solving many important and computationallyintensive tasks in various fields of science and technology. Therefore, it is not surprising that there are several dozen supercomputer centers only in Russia. However, the goals of creating such centers, as well as the range of tasks solved in them, can vary greatly, therefore the structure of supercomputers and the policies for their usage can significantly differ. This leads to the fact that many supercomputer centers live an isolated life – the administrators of such centers tend to solve administration-related tasks on their own, despite the fact that solutions for many similar tasks have already been developed and applied in other centers. This can happen due to different reasons, but in any case, this situation could and should be improved. To do this, it is worth establishing a closer connection between supercomputer centers, which will allow more actively exchanging experience or jointly developing desired system software. In order to understand the current situation in this area, a survey was conducted of representatives among 10 large supercomputer centers in Russia, and its results are presented in this paper. Two relevant topics about using monitoring data in practice and real-life examples of supercomputer functioning improvement are also discussed here in more detail. Their vision on these topics is provided by the system administrators of HSE University, Skoltech and Moscow State University.
HPC systems are complex in architecture and contain millions of components. To ensure reliable operation and efficient output, functioning of most subsystems should be supervised. This is done on the basis of collected data from various logging and monitoring systems. This means that different data sources are used, and accordingly, data analysis can face multiple issues processing this data. Some of the data subsets can be incorrect due to the malfunctioning of used sensors, monitoring system data aggregation errors, etc. This is why it is crucial to preprocess such monitoring data before analyzing it, taking into the consideration the analysis goals. The aim of this paper is, being based on the MSU HPC Center monitoring data, to propose an approach to data preprocessing of HPC monitoring systems, giving some real life examples of issues that may be faced, and recommendations for further analysis of similar datasets. Высокопроизводительные вычислительные системы сложны по архитектуре и содержат миллионы компонент. Чтобы обеспечить надежную работу и эффективную отдачу, необходимо контролировать работу всех их подсистем. Это делается на основе данных, собранных различными системами журналирования и мониторинга. Это означает, что используются разные источники данных, и, соответственно, анализ данных может столкнуться с множеством проблем, связанных с обработкой этих данных. Некоторые из подмножеств данных могут быть неверными из-за неисправности используемых датчиков, ошибок агрегирования данных системы мониторинга и т.д. Вот почему крайне важно проводить предварительную обработку таких данных мониторинга перед их анализом, принимая во внимание цели анализа. Цель этой работы, описать подход к предварительной обработке данных суперкомпьютерных систем мониторинга на основе опыта работы СКЦ МГУ, привести некоторые реальные примеры проблем, с которыми можно при этом столкнуться, а также рекомендации по дальнейшему анализу подобных наборов данных.
Many contemporary HPC systems expose their jobs to substantial amounts of interference, leading to significant run-to-run variation. For example, application runtimes on Theta , a Cray XC40 system at Argonne National Laboratory, vary by up to 70 % , caused by a mix of node-level and system-level effects, including network and file-system congestion in the presence of concurrently running jobs. This makes performance measurements generally irreproducible, heavily complicating performance analysis and modeling. On noisy systems, performance analysts usually have to repeat performance measurements several times and then apply statistics to capture trends. First, this is expensive and, second, extracting trends from a limited series of experiments is far from trivial, as the noise can follow quite irregular patterns. Attempts to learn from performance data how a program would perform under different execution configurations experience serious perturbation, resulting in models that reflect noise rather than intrinsic application behavior. On the other hand, although noise heavily influences execution time and energy consumption, it does not change the computational effort a program performs. Effort metrics that count how many operations a machine executes on behalf of a program, such as floating-point operations, the exchange of MPI messages, or file reads and writes, remain largely unaffected and—rare non-determinism set aside—reproducible. This paper addresses initial stage of an ExtraNoise project, which is aimed at revealing and tackling key questions of system noise influence on HPC applications.
High-performance computing takes a very important place in modern scientific research process. And since all scientists want to solve their problems faster, it is very important to speed up these computations. For these purposes, new algorithms are being developed, new HPC systems appear, etc. However, quite little attention is paid to the efficiency of high-performance computations, which often leads to a vast amount of supercomputer resources being idle. It is vital to change this situation; in particular, it is necessary to show users the importance and necessity of optimizing their applications. One of the main steps in this direction is to help users detect performance issues in their programs, analyze their level of criticality as well as root causes, and eliminate them in order to improve application performance. In this article we describe the research being performed at the Lomonosov Moscow State University aimed at solving this problem. In particular, we analyze the results of supercomputer center users survey, showing their opinion on the efficiency analysis. We also share our vision on the HPC center workflow requirements to support system and applications efficiency analysis. After that, we describe a software tool being developed that allows any supercomputer user to obtain and analyze versatile statistics on performance of his HPC jobs, helping him to detect possible root causes of performance degradation.
There is a variety of known HPC ratings nowadays which represent machine capability for solving a fixed problem, based on a certain algorithm, but these ratings represent a top of the iceberg, and as a rule, one can’t compare application tuning features even for the selected system, and the details of system architecture are not usually described precisely. At the same time lots of efforts are made to describe diverse algorithm features formally, AlgoWiki is one of the most notable recent projects. The idea of Algo500 is joining precise description of computer system with detailed formal descriptions of algorithms using implementation performance data, and building an engine over such joint base to allow various queries, thus giving means of building user-defined ratings regarding selected method, algorithms and/or computer platform features. This paper gives an overview of Algo500 design and some use cases.
The described project is aimed at a complete solution to the problem of joint analysis of the properties of algorithms and features with the architecture of computing systems. This problem arose in the mid-70s of the last century, and over time, its importance in the practice of using computer systems is constantly growing. The main reason is a significant complication of the architecture of computers, which determines a strong dependence of the efficiency of their work on the properties of algorithms and programs. Exactly this dependence leads in practice to a huge gap between real and peak performance indicators, which is typical for all classes of computers from mobile devices to supercomputers of the highest performance range. It is this dependence that leads to a decrease in the quality of work of supercomputer centers and a drop in the efficiency of computer systems below a fraction of a percent. And at the same time, the fundamental nature of the problem itself determines two important facts. First, it is characteristic of all computer systems and centers of the world without exception. Second, practically all scientific groups of the world in all science areas conducting research using high-performance computing systems face this problem.
In order to ensure high performance of a large supercomputer facility, its management must collect and analyze information on many different aspects of this complex operation in a timely manner. One of the most important directions in this area is to analyze the supercomputer workload. In particular, it is necessary to detect the subject areas being the most active in terms of consumed node-hours, the application packages using the GPU with the least efficiency and the reasons for that, the frequency of launching multi-process jobs etc. To do this, it is necessary to collect different types of data from various sources, integrate them within a single solution and develop methods for their analysis. In this article, we describe our approach to solving this problem and show real-life examples of results obtained on the Lomonosov-2 supercomputer.
Managing and administering an HPC center is a real challenge. The Supercomputing Center at Moscow State University is the largest in Russia, with petascale machines running in the interest of thousands of users engaged in hundreds of research projects. Since its early stages of development, the Octoshell system has been instrumental in mastering the equipment and tackling technical issues. The more problems are solved, the more new ideas arise and chances for improvement are discovered. In this paper, we discuss major enhancements recently made to Octoshell, including the integration of new modules (such as extended object description, license management, hardware accounting, and others), and outline plans for the near future.
Every HPC Center holder wants his systems to be productive, and every supercomputer user tries to achieve his research goals as fast as possible, thus wanting his jobs to run efficiently. There are lots of possible approaches for solving such problems. This paper shares the experience of tackling these challenges at the supercomputer center of Moscow State University, the largest HPC center in Russia, hosting Lomonosov-2 system with 4.95 PFlops of peak performance and some smaller systems aimed at providing computing services for hundreds of research projects, running about one thousand jobs per day. The proposed approach covers various levels of role-specific analysis scenarios, from the overall resource utilization and job state distribution down to the precise efficiency analysis of particular jobs including detection of possible performance issues. This approach is being implemented using modules for the Octoshell HPC center management system and used by system administrators and regular users in their daily practice.
Just after the computer era has been started, the Research Computing Center of the Moscow State University was equipped with the most modern computing hardware. These days RCC MSU still operates large scale supercomputers including Lomonosov and Lomonosov-2. Supercomputers are open for research and education society supporting hundreds of projects. The huge numbers of hardware and software components and parameters together with the complexity of architectures implemented raise an extremely important question – how efficiently the supercomputers are used. The efficiency study requires deep monitoring and analysis of all processes inside supercomputers. To improve the efficiency, a set of tools and techniques should be created to make a quick and automated decisions in all sides of supercomputer functioning. The paper is a brief overview of RCC MSU experience in supercomputers productivity improving by use of smart software and analytical techniques.
This paper illustrates the approach for inspecting factors of hardware details and computing problems that impact the performance of typical computing applications and for estimating its influence. The approach extends the idea of supercomputing lists (like TOP500, Graph500, HPCG and others) and combines it with researching algorithms properties (given by AlgoWiki). Possible applications of the approach are described and the main principles of the implementation are given.
Running any computing center is a complex task. With the growth of scales and costs such tasks become challenges. So the top supercomputer sites, being big in everything, have always required special approaches to manage, to control, and to take care of them. At present, large HPC centers can have a variety of totally diverse systems containing up to millions of components, having thousands of users worldwide with the full range of complicated applications. Obviously, tons of data have to be managed in a concerted way to allow such an informational factory functioning. This paper shares the design principles, some implementation details and the roadmap vision regarding the Octoshell HPC center management system, which has been developed and is currently being used in the everyday practice of Moscow State University supercomputer center. This open source system manages Lomonosov and Lomonosov-2 systems with a total of over 5 PFlops peak performance complexes at present, providing multiple tools aimed to tackle most typical workflow tasks both for regular users and system administrators in a single shell.
Osamu Watanabe合作论文数Department of Mathematical and Computing Science
Tokyo Institute of Technology
Japan Chapter of EATCS, the European Association for Theoretical Computer Science.1
Sverker Holmgren合作论文数Division of Scientific Computing
Department of Information Technology
Uppsala University1