In this work, the ability of rare VHE gamma ray selection with neural network methods is investigated in the case when cosmic radiation flux strongly prevails (ratio up to 10^4 over the gamma radiation flux from a point source). This ratio is valid for the Crab Nebula in the TeV energy range, since the Crab is a well-studied source for calibration and test of various methods and installations in gamma astronomy. The part of TAIGA experiment which includes three Imaging Atmospheric Cherenkov Telescopes observes this gamma-source too. Cherenkov telescopes obtain images of Extensive Air Showers. Hillas parameters can be used to analyse images in standard processing method, or images can be processed with convolutional neural networks. In this work we would like to describe the main steps and results obtained in the gamma/hadron separation task from the Crab Nebula with neural network methods. The results obtained are compared with standard processing method applied in the TAIGA collaboration and using Hillas parameter cuts. It is demonstrated that a signal was received at the level of higher than 5.5σ in 21 h of Crab Nebula observations after processing the experimental data with the neural network method.
Imaging atmospheric cherenkov telescopes (IACTs) of the gamma ray observatory TAIGA detect the extesnive air showers (EASs) originating from the cosmic or gamma rays interactions with the atmosphere. Thereby, telescopes obtain images of the EASs. The ability to segregate gamma rays images from the hadronic cosmic ray background is one of the main features of this type of detectors. However, in actual IACT observations, simultaneous observation of the background and the source of gamma rays is needed. This observation mode (called wobbling) modifies images of events, which affects the quality of selection by neural networks. Thus, in this work, the results of the application of neural networks (NN) for the image classification task on Monte Carlo (MC) images of the TAIGA-IACTs are presented. The wobbling mode is considered together with the image adaptation for the adequate analysis by NNs. Simultaneously, we explore several neural network structures that classify events both directly from images or through Hillas parameters extracted from images. In addition, by employing NNs, MC simulation data are used to evaluate the quality of the segregation of rare gamma events with the account of all necessary image modifications.
The TAIGA experimental complex is a hybrid observatory for high-energy gamma-ray astronomy in the range from 10 TeV to several EeV. The complex consists of such installations as TAIGA- IACT, TAIGA-HiSCORE and a number of others. The TAIGA-HiSCORE facility is a set of wide-angle synchronized stations that detect Cherenkov radiation scattered over a large area. TAIGA-HiSCORE data provides an opportunity to reconstruct shower characteristics, such as shower energy, direction of arrival, and axis coordinates. The main idea of the work is to apply convolutional neural networks to analyze HiSCORE events, considering them as images. The distribution of registration times and amplitudes of events recorded by HiSCORE stations is used as input data. The paper presents the results of using convolutional neural networks to determine the characteristics of air showers. It is shown that even a simple model of convolutional neural network provides the accuracy of recovering EAS parameters comparable to the traditional method. Preliminary results of air shower parameters reconstruction obtained in a real experiment and their comparison with the results of traditional analysis are presented.
In recent years, machine learning techniques have seen huge adoption in astronomy applications. In this work, we discuss the generation of realistic synthetic images of gamma-ray events, similar to those captured by imaging atmospheric Cherenkov telescopes (IACTs), using the generative model called a conditional generative adversarial network (cGAN). The significant advantage of the cGAN technique is the much faster generation of new images compared to standard Monte Carlo simulations. However, to use cGAN-generated images in a real IACT experiment, we need to ensure that these images are statistically indistinguishable from those generated by the Monte Carlo method. In this work, we present the results of a study comparing the parameters of cGAN-generated image samples with the parameters of image samples obtained using Monte Carlo simulation. The comparison is made using the so-called Hillas parameters, which constitute a set of geometric features of the event image widely employed in gamma-ray astronomy. Our study demonstrates that the key point lies in the proper preparation of the training set for the neural network. A properly trained cGAN not only excels at generating individual images but also accurately reproduces the Hillas parameters for the entire sample of generated images. As a result, machine learning simulations are a compelling alternative to time-consuming Monte Carlo simulations, offering the speed required to meet the growing demand for synthetic images in IACT experiments.
Imaging atmospheric Cherenkov telescopes are used to record images of extensive area showers caused by high-energy particles colliding with the upper atmosphere. The images are analyzed to determine events’ physical parameters, such as the type and the energy of the primary particles. The distributions of some of the physical parameters can be used as well, for example, to determine the properties of a gamma ray source. The key problem of any experiment is the calibration of experimental data. For this purpose, Monte Carlo simulated data with known values of the physical parameters are used. The main disadvantage of this method is its extremely high requirements for computing resources and the large amount of time spent on modelling. In this paper, we use an alternative approach: Cherenkov telescope images are simulated with conditional variational autoencoders. We compare the characteristics of both the individual images and their Hillas parameter distributions with those of the images generated by the Monte Carlo method.
Generative adversarial networks are a promising tool for image generation in the astronomy domain. Of particular interest are conditional generative adversarial networks (cGANs), which allow you to divide images into several classes according to the value of some property of the image, and then specify the required class when generating new images. In the case of images from Imaging Atmospheric Cherenkov Telescopes (IACTs), an important property is the total brightness of all image pixels (image size), which is in direct correlation with the energy of primary particles. We used a cGAN technique to generate images similar to whose obtained in the TAIGA-IACT experiment. As a training set, we used a set of two-dimensional images generated using the TAIGA Monte Carlo simulation software. We artificiallly divided the training set into 10 classes, sorting images by size and defining the boundaries of the classes so that the same number of images fall into each class. These classes were used while training our network. The paper shows that for each class, the size distribution of the generated images is close to normal with the mean value located approximately in the middle of the corresponding class. We also show that for the generated images, the total image size distribution obtained by summing the distributions over all classes is close to the original distribution of the training set. The results obtained will be useful for more accurate generation of realistic synthetic images similar to the ones taken by IACTs.
High-energy particles hitting the upper atmosphere of the Earth produce extensive air showers that can be detected from the ground level using imaging atmospheric Cherenkov telescopes. The images recorded by Cherenkov telescopes can be analyzed to separate gamma-ray events from the background hadron events. Many of the methods of analysis require simulation of massive amounts of events and the corresponding images by the Monte Carlo method. However, Monte Carlo simulation is computationally expensive. The data simulated by the Monte Carlo method can be augmented by images generated using faster machine learning methods such as generative adversarial networks or conditional variational autoencoders. We use a conditional variational autoencoder to generate images of gamma events from a Cherenkov telescope of the TAIGA experiment. The variational autoencoder is trained on a set of Monte Carlo events with the image size, or the sum of the amplitudes of the pixels, used as the conditional parameter. We used the trained variational autoencoder to generate new images with the same distribution of the conditional parameter as the size distribution of the Monte Carlo-simulated images of gamma events. The generated images are similar to the Monte Carlo images: a classifier neural network trained on gamma and proton events assigns them the average gamma score 0.984, with less than 3 the same time, the sizes of the generated images do not match the conditional parameter used in their generation, with the average error 0.33.
Extensive air showers created by high-energy particles interacting with the Earth atmosphere can be detected using imaging atmospheric Cherenkov telescopes (IACTs). The IACT images can be analyzed to distinguish between the events caused by gamma rays and by hadrons and to infer the parameters of the event such as the energy of the primary particle. We use convolutional neural networks (CNNs) to analyze Monte Carlo-simulated images from the telescopes of the TAIGA experiment. The analysis includes selection of the images corresponding to the showers caused by gamma rays and estimating the energy of the gamma rays. We compare performance of the CNNs using images from a single telescope and the CNNs using images from two telescopes as inputs.
A data life cycle (DLC) is a high-level data processing pipeline that involves data acquisition, event reconstruction, data analysis, publication, archiving, and sharing. For astroparticle physics a DLC is particularly important due to the geographical and content diversity of the research field. A dedicated and experiment spanning analysis and data centre would ensure that multi-messenger analyses can be carried out using state-of-the-art methods. The German-Russian Astroparticle Data Life Cycle Initiative (GRADLCI) is a joint project of the KASCADE-Grande and TAIGA collaborations, aimed at developing a concept and creating a DLC prototype that takes into account the data processing features specific for the research field. An open science system based on the KASCADE Cosmic Ray Data Centre (KCDC), which is a web-based platform to provide the astroparticle physics data for the general public, must also include effective methods for distributed data storage algorithms and techniques to allow the community to perform simulations and analyses with sophisticated machine learning methods. The aim is to achieve more efficient analyses of the data collected in different, globally dispersed observatories, as well as a modern education to Big Data Scientist in the synergy between basic research and the information society. The contribution covers the status and future plans of the initiative.
In this paper, we present the results of comparing container virtualization tools to solve the problem of using idle resources of a supercomputer. On average, as much as 10% of computational resources of a supercomputer may be underloaded due to various reasons. Our basic idea is to maintain an additional queue of low-priority non-parallel jobs that will run on idle resources until a regular job from the main queue of the supercomputer arrives. Upon arrival of the regular job, the low-priority jobs temporarily interrupt their execution and wait for the appearance of new idle nodes to be resumed there. This approach can be implemented by running low-priority jobs in containers and using the container migration mechanism to freeze these jobs and then run them from the point they were frozen at. Thus, the selection of a specific container virtualization tool that is best suited to our goal is an important task. Preliminary analysis allowed us to choose Docker and LXC software products. In this work, we make a detailed comparison of these tools and show why Docker is preferable for solving the above problem.
We propose an approach to utilize idle computational resources of supercomputers. The idea is to maintain an additional queue of low-priority non-parallel jobs and execute them in containers, using container migration tools to break the execution down into separate intervals. We propose a container management system that can maintain this queue and interact with the supercomputer scheduler. We conducted a series of experiments simulating supercomputer scheduler and the proposed system. The experiments demonstrate that the proposed system increases the effective utilization of supercomputer resources under most of the conditions, in some cases significantly improving the performance.
Andreas Haungs∗ 1, Igor Bychkov 2,3, Julia Dubenskaya 4, Oleg Fedorov 5, Andreas Heiss 6, Donghwa Kang 1, Yulia Kazarina 5, Elena Korosteleva 4, Dmitriy Kostunin 1,7, Alexander Kryukov 4, Andrey Mikhailov 2, Minh-Duc Nguyen 4, Frank Polgart 1, Stanislav Polyakov 4, Evgeny Postnikov 4, Alexey Shigarov 2,3, Dmitry Shipilov 5, Achim Streit 6, Victoria Tokareva 1, Doris Wochele 1, Jürgen Wochele 1, Dmitry Zhurov 5 1 Karlsruhe Institute of Technology, IKP, 76021 Karlsruhe, Germany 2 Matrosov Inst. f. System Dynamics and Control Theory, Irkutsk 664033, Russia 3 Irkutsk State University, Irkutsk 664003, Russia 4 Lomonosov Moscow State University, SINP, Moscow 119991, Russia 5 Irkutsk State University, Applied Physics Institute, Irkutsk 664003, Russia 6 Karlsruhe Institute of Technology, SCC, 76021 Karlsruhe, Germany 7 DESY, 15738 Zeuthen, Germany
Deep learning techniques, namely convolutional neural networks (CNN), have previously been adapted to select gamma-ray events in the TAIGA experiment, having achieved a good quality of selection as compared with the conventional Hillas approach. Another important task for the TAIGA data analysis was also solved with CNN: gamma-ray energy estimation showed some improvement in comparison with the conventional method based on the Hillas analysis. Furthermore, our software was completely redeveloped for the graphics processing unit (GPU), which led to significantly faster calculations in both of these tasks. All the results have been obtained with the simulated data of TAIGA Monte Carlo software; their experimental confirmation is envisaged for the near future.
In this work we compare two open source machine learning libraries, PyTorch and TensorFlow, as software platforms for rejecting hadron background events detected by imaging air Cherenkov telescopes (IACTs). Monte Carlo simulation for the TAIGA-IACT telescope is used to estimate background rejection quality. A wide variety of neural network algorithms provided by both libraries can easily be tested on various types of data, which is useful for various imaging air Cherenkov experiments. The work is a component of the Astroparticle.online project, which collaborates with the TAIGA and KASCADE experiments and welcomes any astroparticle experiment to join.
We present a simple set of command line interface tools called Docker Container Manager (DCM) that allow users to create and manage Docker containers with preconfigured SSH access while keeping the users isolated from each other and restricting their access to the Docker features that could potentially disrupt the work of the server. Users can access DCM server via SSH and are automatically redirected to DCM interface tool. From there, they can create new containers, stop, restart, pause, unpause, and remove containers and view the status of the existing containers. By default, the containers are also accessible via SSH using the same private key(s) but through different server ports. Additional publicly available ports can be mapped to the respective ports of a container, allowing for some network services to be run within it. The containers are started from read-only filesystem images. Some initial images must be provided by the DCM server administrators, and after containers are configured to meet one's needs, the changes can be saved as new images. Users can see the available images and remove their own images. DCM server administrators are provided with commands to create and delete users. All commands were implemented as Python scripts. The tools allow to deploy and debug medium-sized distributed systems for simulation in different fields on one or several local computers.
Provenance metadata (PMD) contain key information that is necessary to determine the origin, authorship and quality of relevant data, their storage and usage consistency, and for interpretation and confirmation of relevant scientific results. The need for PMD is especially important when Big Data are jointly processed by several research teams, which is a very common practice in many scientific areas of late. Although a number of projects have been implemented in recent years to create management systems for such metadata, the vast majority of the implemented solutions are centralized, which is poorly suited to current trends of working in distributed environments and using metadata by organizationally unrelated or loosely coupled communities of researchers. We propose to solve this problem by employing a new approach to creating a distributed registry of provenance metadata based on blockchain technology and smart contracts. We have investigated the problem of the optimal choice of the type of blockchain for such a system, as well as the optimal choice of the blockchain platform. The architecture and algorithms of the system operation, as well as its interaction with the distributed storage resources management systems, are proposed.
In the frame of the Karlsruhe-Russian Astroparticle Data Life Cycle Initiative it was proposed to deploy an educational resource astroparticle.online for the training of students in the field of astroparticle physics. This resource is based on HUBzero, which is an open-source software platform for building powerful websites, which supports scientific discovery, learning, and collaboration. HUBzero has been deployed on the servers of Matrosov Institute for System Dynamics and Control Theory. The educational resource astroparticle.online is being filled with the information covering cosmic messengers, astroparticle physics experiments and educational courses and schools on astroparticle physics. Furthermore, the educational resource astroparticle.online can be used for online collaboration. We present the current status of this project and our first experience of application of this service as a collaboration framework.
Modern large-scale astroparticle setups measure high-energy particles, gamma rays, neutrinos, radio waves, and the recently discovered gravitational waves. Ongoing and future experiments are located worldwide. The data acquired have different formats, storage concepts, and publication policies. Such differences are a crucial point in the era of Big Data and of multi-messenger analysis in astroparticle physics. We propose an open science web platform called ASTROPARTICLE.ONLINE which enables us to publish, store, search, select, and analyze astroparticle data. In the first stage of the project, the following components of a full data life cycle concept are under development: describing, storing, and reusing astroparticle data; software to perform multi-messenger analysis using deep learning; and outreach for students, post-graduate students, and others who are interested in astroparticle physics. Here we describe the concepts of the web platform and the first obtained results, including the meta data structure for astroparticle data, data analysis by using convolution neural networks, description of the binary data, and the outreach platform for those interested in astroparticle physics. The KASCADE-Grande and TAIGA cosmic-ray experiments were chosen as pilot examples.
Modern detectors of cosmic gamma-rays are a special type of imaging telescopes (air Cherenkov telescopes) supplied with cameras with a relatively large number of photomultiplier-based pixels. For example, the camera of the TAIGA-IACT telescope has 560 pixels of hexagonal structure. Images in such cameras can be analysed by deep learning techniques to extract numerous physical and geometrical parameters and/or for incoming particle identification. The most powerful deep learning technique for image analysis, the so-called convolutional neural network (CNN), was implemented in this study. Two open source libraries for machine learning, PyTorch and TensorFlow, were tested as possible software platforms for particle identification in imaging air Cherenkov telescopes. Monte Carlo simulation was performed to analyse images of gamma-rays and background particles (protons) as well as estimate identification accuracy. Further steps of implementation and improvement of this technique are discussed.
Modern supercomputer schedulers on average may leave ~10% and sometimes as much as 30% of the computational resources idle. One possible approach to increase the load is to use an additional queue of low-priority jobs small enough to fit into the schedule gaps. We propose to use this approach for non-parallel jobs with arbitrary runtime wrapped in containers to allow them to be saved and migrated to other nodes or back to the queue. As a result, all the idle nodes can be used for computations. We also estimate the increase in average load and utilization efficiency that can be achieved using this approach.
A Streit合作论文数J??lich Supercomputing Centre3
Sergei A. Abramov合作论文数Russian Academy of Sciences;Dorodnicyn Computing Centre1