The ATLAS experiment at the Large Hadron Collider has a complex heterogeneous distributed computing infrastructure, which is used to process and analyse exabytes of data. Metadata are collected and stored at all stages of data processing and physics analysis. All metadata could be divided into operational metadata to be used for the quasi on-line monitoring, and archival to study the behaviour of corresponding systems over a given period of time (i.e. long-term data analysis). Ensuring the stability and efficiency of complex and large-scale systems, such as those in the ATLAS Computing, requires sophisticated monitoring tools, and the long-term monitoring data analysis becomes as important as the monitoring itself. Archival metadata, which contains a lot of metrics (hardware and software environment descriptions, network states, application parameters, errors) accumulated for more than a decade, can be successfully processed by various machine learning (ML) algorithms for classification, clustering and dimensionality reduction. However, the ML data analysis, despite the massive use, is not without shortcomings: the underlying algorithms are usually treated as "black boxes", as there are no effective techniques for understanding their internal mechanisms. As a result, the data analysis suffers from the lack of human supervision. Moreover, sometimes the conclusions made by algorithms may not be making sense with regard to the real data model. In this work we will demonstrate how the interactive data visualization can be applied to extend the routine ML data analysis methods. Visualization allows an active use of human spatial thinking to identify new tendencies and patterns found in the collected data, avoiding the necessity of struggling with the instrumental analytics tools. The architecture and the corresponding prototype of Interactive Visual Explorer (InVEx) - visual analytics toolkit for the multidimensional data analysis of ATLAS computing metadata will be presented. The web-application part of the prototype provides an interactive visual clusterization of ATLAS computing jobs, search for computing jobs non-trivial behaviour and its possible reasons.
The Interactive Visual Explorer (InVEx) application is designed as a visual analytics tool for Big Data analysis. Visual analytics is an integral approach to data analysis, combining methods of intellectual data analysis with advanced interactive visualization. One of the main objectives of InVExis to process large data samples by decreasing their level of detail (LoD).The proposed approach includes clustering as well as flexible grouping by different parameters, providing the exploration of data from the lowest to the highest level of details. The results of grouping and clusterization arevisualized using interactive 3D scene and parallel coordinates, allowing the user to gain insight into data, to explore hidden correlations and trends of parameters.
The effect of the carbon fibers coupling layer on the occurrence of the triple-shape memory effect of polyurethane reinforced is studied. Using thermomodulated differential scanning calorimetry, structural changes in a sample of Carbon fiber reinforced polyurethane with coupling layer were determined. The influence of the diffusion adhesion mechanism on the thermomechanical characteristics of the triple-shape memory effect of the polyurethane composite material is established.
The development of the Interactive Visual Explorer (InVEx), a visual analytics tool for the computing metadata of the ATLAS experiment at LHC, includes research of various approaches for data handling both on server and client sides. InVEx is implemented as a web-based application which aims at the enhancing of analytical and visualization capabilities of the existing monitoring tools and facilitates the process of data analysis with the interactivity and human supervision. The current work is focused on the architecture enhancements of the InVEx application. First, we will describe the user-manageable data preparation stage for cluster analysis. Then, the Level-of-Detail approach for the interactive visual analysis will be presented. It starts with the low detailing, when all data records are grouped (by clustering algorithms or by categories) and aggregated. We provide users with means to look deeply into this data, incrementally increasing the level of detail. Finally, we demonstrate the development of data storage backend for InVEx, which is adapted for the Level-of-Detail method to keep all stages of data derivation sequence.
The article presents a design of a precision large-size antenna reflector that made of polymer composite materials (PCM) based on carbon fibers. Such reflectors are used for operation in high frequency ranges, since they have a low coefficient of linear thermal expansion and a high modulus of elasticity. Therefore, the main task of this work is to design the geometric accuracy of the reflector working surface from composite materials with a diameter of more than 10 meters and a frequency range of 42.5-45.5 GHz. The developed model of the reflector includes a power frame, segments of the reflecting surface and a hub. The power frame of the reflector consists of flat trusses supplemented with rods, so that during assembly a spatial construction with axial symmetry is formed. Segments are three-layer casings of polymer composite materials with filler. The proposed model of the reflector was analyzed using the finite element method with boundary conditions: a wind load of 20 m/s in the opening of the reflector; impact of gravity on the reflector, oriented to the zenith. The wind load was modeled as a uniformly distributed pressure applied to the segments. The obtained mean square deviation (SDE) of the geometry and natural oscillation frequency of a reflector made from polymer composite materials based on carbon fibers is sufficient for operation of the satellite earth station in high radio-frequency ranges Ka, Q and V.
Interactive visual analysis tools bring the ability of the real-time discovery of knowledge in large and complex datasets using visual analytics. It involves multiple iterations of data processing using various data handling approaches and the efficiency of the whole chain of the analysis process depends on the performance of chosen techniques and related implementations, as well as the quality of applied methods. Stages, where data processing includes intellectual handling (i.e., data mining and machine learning), which are the most resource-intensive, require a distinct attention for evaluation of different approaches. Clustering is one such machine learning technique that is commonly used to discover groups of data objects for further analysis. This work is focused on evaluation of clustering algorithms within the interactive visual analysis toolkit InVEx (Interactive Visual Explorer). InVEx represents a visual analytics approach aimed at cluster analysis and in-depth study of implicit correlations between multidimensional data objects. It is originally designed to enhance the analysis of computing metadata of the ATLAS experiment at the LHC for operational needs, but it also provides the same capabilities for other domains to analyze large amounts of multidimensional data. The experiments and evaluation processes are carried out using operational data from the supercomputer at the Lomonosov Moscow State University. These processes include benchmark tests to assess the relative performance between chosen clustering algorithms and corresponding metrics to assess the quality of produced clusters. Obtained results will be used as guidelines in assisting users in a process of visual analysis using InVEx.
The ATLAS experiment at the LHC processes, analyses and stores vast amounts of data, which is either recorded by the detector or simulated worldwide using Monte Carlo methods. ATLAS Computing metadata is generated at very high rates and volumes. The necessity to analyze this metadata is constantly increasing, since the heterogeneous, distributed and dynamically changing computing infrastructure requires sophisticated optimization decisions, made by human or/and by machines. Visual analytics is one of the methods facilitating the analysis of massive amounts of data (structured, semi-structured, and unstructured) which leverages human judgement by means of interactive visual representations. Given the huge number of ATLAS computing jobs that need to be visualized simultaneously for error investigations or other optimization processes, resources of the client application responsible for such visualization may reach its limits. Data objects that share similar feature values can be represented and visualized as a single group, thus initial large data sample would be represented at different levels of detail. This approach will also avoid client overload. In this paper we evaluate implementations of k-means-based Level-of-Detail generator method applied to the metadata of ATLAS jobs. This method is used in the visual analytics application InVEx (Interactive Visual Explorer) that is under development, and which is based on 3-dimensional interactive visualization of multidimensional data.
This article presents a research on Hexagonal Boron Nitride (h-BN) monolayer cell strain effect 2 % and 4 %. Structure of h-BN with nitrogen vacancy, with boron vacancy and with divacancy was considered for this. The calculations were carried out within framework of the density functional formalism with gradient corrections and using the VASP package. Vanderbilt Ultra-Soft Pseudopotential was used in the course of the calculations. It is possible to conclude that nitrogen vacancies are the most stable, regardless of monolayer deformation on the results obtained. Understanding of atomic scale stability and dynamics of defects in such systems is crucial for predicting their properties and applications in electronics.
Scientific computing has advanced in the ways it deals with massive amounts of data, since the production capacities have increased significantly for the last decades. Most large science experiments require vast computing and data storage resources in order to provide results or predictions based on the data obtained. For scientific distributed computing systems with hundreds of petabytes of data and thousands of users it is important to keep track not just of how data is distributed in the system, but also of individual users’ interests in the distributed data (reveal implicit interconnection between user and data objects). This however requires the collection and use of specific statistics such as correlations between data distribution, the mechanics of data distribution, and mainly user preferences. This work focuses on user activities (specifically, data usages) and interests in such a distributed computing system, namely PanDA (Production ANd Distributed Analysis system). PanDA is a high-performance workload management system originally designed to meet production and analysis requirements for a data-driven workload at the Large Hadron Collider Computing Grid for the ATLAS Experiment hosted at CERN (the European Organization for Nuclear Research). In this work we are going to investigate whether data collection that was gathered in the past in PanDA shows any trends indicating that users could have mutual interests that would be kept for the next data usages (i.e., data usage patterns), using data mining techniques such as association analysis, sequential pattern mining, and basics of the recommender system approach. We will show that such common interests between users indeed exist and thus could be used to provide recommendations (in terms of the collaborative filtering) to help users with their data selection process.
One of the most important aspects in any computing distribution system is efficient data replication over storage or computing centers, that guarantees high data availability and low cost for resource utilization. In this paper we propose a data distribution scheme for the production and distributed analysis system PanDA at the ATLAS experiment. Our proposed scheme is based on the investigation of data usage. Thus, the paper is focused on the main concepts of data popularity in the PanDA system and their utilization. Data popularity is represented as the set of parameters that are used to predict the future data state in terms of popularity levels.
The workflow management process should be under control of a specific service that is able to forecast the processing time dynamically according to the status of the processing environment and workflow itself, and to react immediately on any abnormal behavior of the execution process. Such situational awareness analytic service would provide the possibility to monitor the execution process, to detect the source of any malfunction, and to optimize the management process. The stated service for the second generation of the ATLAS Production System (ProdSys2, an automated scheduling system) is based on predictive analytics approach. Its primary goal is to estimate the duration of the data processings (in terms of ProdSys2, it is task and chain of tasks) with possibility for later usage in decision making processes. Machine learning ensemble methods are chosen to estimate completion time (i.e., “Time-To-Complete”, TTC) for every (production) task and chain of tasks, and “abnormal” task processing times would warn about possible failure state of the system. This is the primary phase of the service development that also includes the strategy for its precision enhancement. The first implementation of such analytic service is designed around Task TTC Estimator tool and it provides a comprehensive set of options to adjust the analysis process and possibility to extend its functionality.
The second generation of the Production System (ProdSys2) of the ATLAS experiment (LHC, CERN), in conjunction with the workload management system PanDA (Production and Distributed Analysis), represents a complex set of computing components that are responsible for defining, organizing, scheduling, starting and executing payloads in a distributed computing infrastructure. ProdSys2/PanDA are responsible for all stages of (re)processing, analysis and modeling of raw and derived data, as well as simulation of physical processes and functioning of the detector using Monte Carlo methods. The prototype of the ProdSys2 Predictive Analytics (P2PA) service is an essential part of the growing analytical service for the ProdSys2 and it will play a key role in the ATLAS distributed computing. P2PA uses such tools as Time-To-Complete (TTC) estimation towards units of the processing (i.e., tasks, chains and groups of tasks) to control the processing state and rate, and to be able to highlight abnormal operations and executions (e.g., to discover stalled processes). It uses methods and techniques of machine learning to obtain corresponding predictive models and metrics that are aimed to characterize the current system's state and its changes over a short period of time.
Gergely V. Zaruba合作论文数Department of Computer Science and Engineering;University of Texas at Arlington2