Nowadays, Artificial Intelligence (AI) and Machine Learning (ML) are being used as critical technologies regardless of fields. For network security monitoring, AI/ML techniques also broadly adopted to detect various threats from enormous network traffic. It definitely can enable us to make quicker and more accurate threats detection; however, since AI systems are often operated as black boxes, it is hard to understand the cause of detection without additional analysis. To solve these issues, there have been several attempts using explainable-AI (XAI) techniques to interpret decision processes of AI. Nevertheless, it still has limitations especially in terms of response at real-world Security Operations Centers (SOCs), because it only focuses on revealing an importance of pre-defined features. This paper proposes a practical and security monitoring-response friendly XAI framework, which improves rapid and accurate decision-making by human agents of SOCs. Particularly, the framework provides an intuitive evidence of cyber attacks or anomalies through four steps as follows: 1)extracting semantic features, 2)scoring threats importance, 3)ranking risk influences and 4)visualizing decision evidences. For evaluation purposes, we work with a dataset consisting of real-world network traffic and annotated threats factors by security monitoring experts. Furthermore, experimental results demonstrate the effectiveness of our proposed framework in terms of a decision-making support for real-time security monitoring.
With the development of computer networks, the amount of network traffic is explosively increasing. In addition, the importance of cyber security is being highlighted as cyber threats increase accordingly. In general, rule-based detection approaches have been used to detect cyber threats. The detection rules used in these are broadly set up to reliably detect cyber threats, resulting in too many unnecessary events. This leads to unanalyzed events, which can lead to severe security incidents. To solve this problem, recently, researches on AI-based cyber threat detection system that learns network traffic information and automatically generates detection rules are being conducted. Most of them have used complex model with sophisticated structures or feature engineering techniques so that AI models can learn as much information as possible. But, these are difficult to use in real-world security monitoring environment where quick decisions need to be made in real time, and are not suitable for that environments because they have been trained and verified through only open datasets. In this paper, we propose an AI-based cyber threat detection system that efficiently learns security event characteristics without any complicated process using tree-based model which efficient to learning tabular data. The proposed system detects cyber threats by learning security event characteristics using only information provided from security devices without complicated feature extraction process. In addition, rather than using the used information as a simple value, the value is transformed through a simple process so that the model can learn the event characteristics more effectively. Using the simplicity of the proposed method, it is expected that it can be applied to the real-world environments, and the possibility of this is demonstrated through real-world data.
With the development of networks and the increase in the number of network devices, the number of cyber attacks targeting them is also increasing. Since these cyber-attacks aim to steal important information and destroy systems, it is necessary to minimize social and economic damage through early detection and rapid response. Many studies using machine learning (ML) and artificial intelligence (AI) have been conducted, among which payload learning is one of the most intuitive and effective methods to detect malicious behavior. In this study, we propose a preprocessing method to maximize the performance of the model when learning the payload in term units. The proposed method constructs a high-quality learning data set by eliminating unnecessary noise (stopwords) and preserving important features in consideration of the machine language and natural language characteristics of the packet payload. Our method consists of three steps: Preserving significant special characters, Generating a stopword list, and Class label refinement. By processing packets of various and complex structures based on these three processes, it is possible to make high-quality training data that can be helpful to build high-performance ML/AI models for security monitoring. We prove the effectiveness of the proposed method by comparing the performance of the AI model to which the proposed method is applied and not. Forthermore, by evaluating the performance of the AI model applied proposed method in the real-world Security Operating Center (SOC) environment with live network traffic, we demonstrate the applicability of the our method to the real environment.
Analyzing social media has become a common way for capturing and understanding people's opinions, sentiments, interests, and reactions to ongoing events. Social media has thus become a rich and real-time source for various kinds of public opinion and sentiment studies. According to psychology and neuroscience, human emotions are known to be strongly dependent on sensory perceptions. Although sensation is the most fundamental antecedent of human emotions, prior works have not looked into their relation to emotions based on social media texts. In this paper, we report the results of our study on sensation effects that underlie human emotions as revealed in social media. We focus on the key five types of sensations: sight, hearing, touch, smell, and taste. We first establish a correlation between emotion and sensation in terms of linguistic expressions. Then, in the second part of the paper, we define novel features useful for extracting sensation information from social media. Finally, we design a method to classify texts into ones associated with different types of sensations. The sensation dataset resulting from this research is opened to the public to foster further studies.
As an essential system for protecting internal networks and valuable information, the firewall monitors and controls network traffic in terms of access control, authentication, logging, and auditing. In particular, it carries out both allowing and blocking communications between internal and external networks based on proper Access Control List (ACL). However, a complex ACL along with huge network environments lead to exposing vulnerabilities and communication problems, because of anomalies among policies. Even though various techniques and applications combined with visualization approaches have been proposed, there is still a lack of usability caused by not only the limitation of the text-based interface but also the complexity of practical use. In order to solve these problems, this work proposes a 3D-based hierarchical visualization method, namely F/Wvis, for intuitive ACL management and analysis. The F/Wvis, particularly, supports ACL management for a large-scale network as well as analysis of detail anomalies on policies by providing a drill-down user interface through the hierarchical visualization approach. Further, the implemented system is evaluated against popular tools by network security experts to identify the usability and effectiveness in real-world situations (a demonstration video is available at: https://bit.ly/34ooEDc).
With a paradigm shift to untact environments, security threats on the network also have been significantly increasing all over the world. To monitor and detect intrusion attempts under enormous network traffic, Security Operation Center (SOC) essentially exploits various security devices. Above all, Network Intrusion Detection System (NIDS) has been operated in public/private sectors as a spearhead to fight against cyber threats. In particular, state-of-the-art technologies, especially ML and AI, have been being studied to achieve quick and accurate intrusion detection. Despite much effort to guarantee a secure network, however, SOCs are still struggling for overcoming various types of threats as well as attacks of similar form with benign traffic. Even though the advanced techniques may find out a complex and unknown attack, operating and managing them in real-world situations cause counterproductively more pressure to agents in the SOC. In order to solve these difficulties, this study introduces an easy-to-use framework to build intrusion detection models based on AI techniques, as well as to operate them depending on a situation using a graphical user interface. The framework supports generating various types of AI- and ML-based intrusion detection models with optimized parameters by only a few steps. Furthermore, an interactive graphical interface makes it easier to manage detection models according to different threat situations. Finally, the performance of models made by the framework is evaluated in terms of accuracy, especially under the real-world SOC environment with live network traffic.
Analyzing social media has become a common way for capturing and understanding people's opinions, sentiments, interests, and reactions to ongoing events. Social media has thus become a rich and real‐time source for various kinds of public opinion and sentiment studies. According to psychology and neuroscience, human emotions are known to be strongly dependent on sensory perceptions. Although sensation is the most fundamental antecedent of human emotions, prior works have not looked into their relation to emotions based on social media texts. In this paper, we report the results of our study on sensation effects that underlie human emotions as revealed in social media. We focus on the key five types of sensations: sight, hearing, touch, smell, and taste. We first establish a correlation between emotion and sensation in terms of linguistic expressions. Then, in the second part of the paper, we define novel features useful for extracting sensation information from social media. Finally, we design a method to classify texts into ones associated with different types of sensations. The sensation dataset resulting from this research is opened to the public to foster further studies.
Point clouds have become a primitive and fundamental material for manifold spatial representations. It can precisely render real-world environments as high-density points which include three-dimensional (3D) coordinates (x, y & z) and other features (color, intensity, and so on). Accordingly, various applications, including robot navigation and self-driving, make use of point clouds not only to detect near objects but to comprehend overall geospatial surroundings. However, it is challenging to exploit the point clouds in terms of spatial query processing in traditional database systems because of its enormous volume and nonstructural formats. In this paper, we propose an efficient method for the manipulation of 3D point cloud based on a Discrete Global Grid System (DGGS). As DGGS represents the Earth as hierarchical sequences of equal area/volume tessellations, it provides an accurate partitioning to integrate and analyze big geospatial data, unlike a base64 geohash representation. This study extends our previous DGGS-based encoding/decoding work to process 3D range queries with more than 64 bits for precise 3D coordinates of point clouds. In particular, we apply PH-tree as a multi-resolution tessellation storage and indexing structure for 3D bounding box queries. The experimental results show that our query processing significantly outperforms the baseline with a linear quadtree. Also, we present the encoding/decoding efficiency of converting large Morton codes from geographic coordinates by using the combination of bit interleaving and lookup tables.
As security threats rapidly spread all over the world, it is critical that network traffic is monitored and protected from abnormal attacks during 24/7. Even though various security devices (Rep., network intrusion detection system, NIDS) had utilized to guarantee a solid network security, it still depends on human being due to complex patterns from unknown threats. This study introduces a graphical interactive system for representing and understanding multivariate cybersecurity attacks. In particular, the interface enhances intuitive judgments combined with machine learning-based analysis of suspicious traffic
With the development of mobile surveying and mapping technologies, point cloud data has been emerging in a variety of applications including robot navigation, self-driving drones/vehicles, and three-dimensional (3D) urban space modeling. In addition, there is an increasing demand for the database management system to share and reuse point cloud data, unlike being treated as archive files in the traditional uses and applications. However, database scalability needs to be explored to process and manage a massive volume of point cloud data defined by a 3D (X, Y, and Z) coordinates system. The typical approach to handle big data and distribute it across multiple nodes is data partitioning. Geohashing is a popular way to convert a latitude/longitude spatial point into a code/string and has used for storing data into buckets of the grid. Many methods of handling big geospatial data, especially NoSQL databases, are based on the geohashing techniques. In this paper, we propose an efficient method to encode/decode 3D point cloud in a Discrete Global Grid System (DGGS) that represents the Earth as hierarchical sequences of equal area/volume tessellations, similar to geohash. The current geohash of base36 has the difficulties of working with high-resolution 3D point clouds for data storage, filter, integration, and analytics because of its limitation of cell size and unequal areas. We employ DGGS-based Morton codes with more than 64 bits for precise 3D coordinates of point cloud and compare the encoding/decoding performance between two implementations: using strings and using the combination of bit interleaving and lookup tables.
Sharing data securely and reliably is challenging, mainly when dealing with big spatial data. Data portals, web services, and platforms often struggle when uploading and downloading such data, requiring significant investments in IT infrastructure and expensive high-bandwidth network connectivity to achieve adequate performance for many applications. Dotloom aims to change this and make the sharing of data easier, straightforward, secure, and efficient. Dotloom is a distributed data platform for synchronizing, replicating, indexing, and processing terabytes of point-cloud data with peer-to-peer technologies. The distributed nature allows instant exchange between data producers and data consumers. Processing pipelines have the power to stream data from multiple peers, and the generated output can be shared again instantly. Remote indexing can be implemented using partial data, reducing transfer costs. Building on the existing "DAT Project" infrastructure, Dotloom adds the functionality needed to manage, query, and visualize point-cloud data. These novel features of Dotloom have the potential to not only transform how we deal with point-clouds but to be influential across the wider big-data research and development community.
Recently, IoT (Internet of Things) technology is applied in various fields to increase convenience and usability. In this paper, we propose a mobile atmospheric pollution monitoring system as an IoT application service. The proposed system consists of a measurement part and a monitoring part. The measurement part collects the concentration of fine dust and the monitoring part displays the collected data utilizing graph and map. This paper discusses design and prototype implementation of the proposed system.
With the increase in the use of 3D scanner to sample the earth surface, there is a surge in the availability of 3D spatial data. 3D spatial data contains a wealth of information and can be of potential use if integrated, processed and analyzed in real-time. The 3D spatial data is generated as continuous data stream, however due to its size, velocity and inherent noise, it is processed offline. Many applications require real-time processing and analysis of spatial stream, for-instance, forest fire management, real-time road traffic analysis, disaster engulfed areas monitoring, etc., however they suffer from slow offline processing of traditional systems. This paper presents and demonstrates a robust and scalable pipeline for the real-time processing and analysis of 3D spatial streams. An experimental evaluation is also presented to prove the effectiveness of the proposed framework.
With the development of social network services, various phenomena can be shared easily and rapidly through human natural language, including not only natural, but also social-cultural phenomena. Consequently, analyses of social media have appreciated in value for understanding human behaviors to grasp public interests or sentiments, as both the medium and outcome of human experiences. From the state of the art psychology and neuroscience, human behaviors, regarding both physical and linguistic aspects, are mostly dependent on sensory perceptions under the realm of the subconscious. Even though sensation is the most fundamental element to understand human behaviors, the rack of background resources make it hard to study the social sensation comparing with the sentimental or opinion mining. This paper focuses on building sensation knowledges to obtain useful human perceptual experiences in natural language expressions, as a requisite for the social sensation analysis. We try to approach the constructing lexicons as a sensation knowledge from two viewpoints, such as a deep learning and lexicon based methods. Then we classify social media text based on the lexicons with considering a part of speech as well as semantic meanings of each word. Finally, we identify which knowledge has a good performance to distinguish sensation expressions from social media data in terms of accuracy and and F-score.
Movements of urban citizens largely take part in geo-social urban dynamics, since they exploit diverse urban districts as necessary. However, it is a non-trivial task to measure city-wide crowd movements and analyze them to understand how we exploit a city space for our daily lives. In this paper, we attempt to capture and take advantages of urban crowd movements by exploiting taxis as a sensor capable of monitoring city-wide, continuous, natural and crowd-sourced movements indirectly. Significantly, we propose a road-centric data space model, with which a variety of heterogenous sensor data can be represented in a common form along roads, enabling to hide heterogenous types of primitive movement logs and to support for convenient data integration in terms of roads, time and sensed urban phenomena. Based on the road-centric integration of taxi-based crowd movements, we classify urban districts according to latent temporal road utilization patterns extracted by a Non-negative Matrix Factorization method.
Analyses of social media have increased in importance for understanding human behaviors, interests, and opinions. Business intelligence based on social media can reduce the costs of managing customer trend complexities. This paper focuses on analyzing sensation information representing human perceptual experiences in social media through the five senses: sight, hearing, touch, smell, and taste. First a measurement is defined to estimate social sensation intensities, and subsequently sensation characteristics on geo-social media are identified using geo-spatial footprints. Finally, we evaluate the accuracy and F-measure of our approach by comparing with baselines.
Ki-Joune Li合作论文数Department of Computer Science
Pusan National University3
Gil-Jin Jang合作论文数School of ECE, Ulsan National Institute of Science and Technology (UNIST)1