Just after the computer era has been started, the Research Computing Center of the Moscow State University was equipped with the most modern computing hardware. These days RCC MSU still operates large scale supercomputers including Lomonosov and Lomonosov-2. Supercomputers are open for research and education society supporting hundreds of projects. The huge numbers of hardware and software components and parameters together with the complexity of architectures implemented raise an extremely important question – how efficiently the supercomputers are used. The efficiency study requires deep monitoring and analysis of all processes inside supercomputers. To improve the efficiency, a set of tools and techniques should be created to make a quick and automated decisions in all sides of supercomputer functioning. The paper is a brief overview of RCC MSU experience in supercomputers productivity improving by use of smart software and analytical techniques.
Современные вычисления устойчиво ассоциируются с параллелизмом на всех уровнях. От персональных планшетов и телефонов вплоть до лучших суперкомпьютеров отличаются масштабы и подходы к решению задач, но идея распределения вычислительной работы присутствует повсюду. Изучение возможностей параллельных вычислений для портативных платформ началось в Научно-исследовательском вычислительном центре МГУ им. М.В.Ломоносова несколько лет назад, в 2014 году. Были разработаны версии популярного бенчмарка Linpack для Android и iOS. Четыре года доступности рейтинга для пользователей вылились в почти двадцать тысяч результатов тестов. В статье рассматриваются факты и тенденции, полученные в результате анализа этих данных. В том числе, полученные результаты подтверждают целесообразность дальнейшего развития рейтинга мобильных устройств. Contemporary computing is strongly associated with parallelism. From personal tablets and phones and up to the top supercomputers, the scales and approaches differ, but the idea of work distribution is present everywhere. A study of portable platforms parallel computing capabilities was started at the Research Computing Center of Lomonosov Moscow State University several years ago, in 2014. There were Android and iOS implementations of popular benchmark developed. Four years of the test availability for regular users turned into almost 20 thousand of test results. This paper addresses the investigation of facts and trends, mined out of this data, proving the reasons for the further Mobile Linpack development.
At RCC MSU we are working on the system for securing of active control and efficient autonomous operating of supercomputers. This system is been implemented at MSU Supercomputing Center. The paper describes system installation, setting and usage experience for the control of the "Chebyshev" supercomputer.
An approach to implementation of supercomputing center control system based on supercomputer graph model has been proposed in RCC MSU. The Octotron system has been developed on the basis of this approach, which is being tested in MSU Supercomputing Center currently. The article describes challenges and tasks, encountered by authors while developing and running the system on supercomputers «Chebyshev» and «Lomonosov». It also includes overview of graph tools used, brief description of modeling language, model visualization and monitoring data import.
State-of-the-art supercomputer is extremely complex, expensive and energy-saturated system. Every component of supercomputer is unreliable and can fail any time. In RCC MSU we are working on the system aimed to eliminate bad after-effects of hardware and software failures and to secure a reliable and efficient autonomous functioning of supercomputers. The system is based on the supercomputer model represented as multi-graph.
An approach to implementation of supercomputing center control system based on supercomputer graph model has been proposed in RCC MSU. The Octotron system has been developed on the basis of this approach, which is being tested in MSU Supercomputing Center currently. The article describes challenges and tasks, encountered by authors while developing and running the system on supercomputers «Chebyshev» and «Lomonosov». It also includes overview of graph tools used, brief description of modeling language, model visualization and monitoring data import.
The current state of the X-Com metacomputing system developed in RCC MSU is described. The system architecture and technological base has totally been redesigned with regard for the experience of its practical use. A new version of the X-Com system highlights such features as scalability and support of distributed environments with a supercomputer performance level. Buffering server mechanisms, means of file synchronization and network security have also been realized in the system.