Event sequence data consists of discrete events that happen over time. By grouping events based on common entities and ordering them chronologically, they form sequences. Events are registered in different domains, ranging from healthcare to logistics. Collections of these sequences typically represent high-level processes for users to discover, identify, and analyze. This discovery is challenging, given that sequences in real-world scenarios can grow long, have many events, many attribute dimensions of events, and/or various event categories. However, limited research focuses on analyzing long event sequences, the focus of this paper. We present LoLo, an interactive visual analytics method based on the analysis of multi-level structures in long event sequence collections. LoLo introduces a strategy to split the sequence collection into meaningful data-driven stages, where the definition of a stage facilitates interpretation and injection of domain knowledge. The stages have different levels, which represent high-level processes taking into account high-level changes (global staging) combined with local sequence variations (local staging). We demonstrate the effectiveness of LoLo by comparing it to a baseline and present two use cases, one is evaluated with two users and the other by us, on real-world data sets showing that our staging method can capture the semantic content in stages and users appreciate being able to switch between different levels of detail.
Policy comprehension is crucial for ensuring data protection. Yet, policies written in flexible and expressive languages such as XACML are not easy to comprehend. In this work, we propose a visualization framework to facilitate the comprehension of XACML policies and their evaluation. Our framework shows a tree representation of the XACML policies to be enforced and highlights the contribution of its policy elements to the overall access decision, thus supporting the understanding of how this decision resulted from the interplay between possibly conflicting access requirements. We implemented our visualization framework as an extension to SAFAX, an XACML-based framework that offers authorization as a service.
International Revenue Sharing Fraud (IRSF) is the most persistent type of fraud in the telco industry. Hackers try to gain access to an operator's network in order to make expensive unauthorized phone calls on behalf of someone else. This results in massive phone bills that victims have to pay while number owners earn the money. Current anti-fraud solutions enable the detection of IRSF afterwards by detecting deviations in the overall caller's expenses and block phone devices to prevent attack escalation. These solutions suffer from two main drawbacks: (i) they act only when financial damage is done and (ii) they offer no protection against future attacks. In this paper, we demonstrate how unsupervised machine learning can be used to discover fraudulent calls at the moment of their establishment, thereby preventing IRSF from happening. Specifically, we investigate the use of Isolation Forests for the detection of frauds before calls are initiated and compare the results to an existing industrial post-mortem anti-fraud solution.
Automation is a popular and very important topic, but with our brains still outperforming Artificial Intelligence (AI) techniques, humans are indispensable in security. Especially with respect to pattern recognition and contextual reasoning, human are superior in keeping false positive rates of automated techniques to a minimum. We therefore still have a job to do and cannot go to the beach.
2018 is a difficult year to summarize for Infosec. After the initial flurry of activity around Spectre and Meltdown in the beginning of January, we ended the year with global supply chain concerns brought about by the Super Micro story. Throughout the year we saw the geopolitical dilemmas of 2018 manifest in cyber security issues. Technology giants like Facebook and Google had a security reckoning. However in pure scariness the medical data breaches of MyHeritage (DNA) and MyFitnessPal (health) rank higher. The Starwood Marriot Hotel breach made every travelling executive nervous for the rest of the year, but probably not as nervous as the incident of CEO Fraud at Pathe. In an effort to alleviate some of that impact we are proud to publish the 6th European Cyber Security Perspectives (ECSP) report. The 2019 issue is filled with great articles from our partners ranging from government, universities and private companies. Special thanks goes out to all the partners who have submitted an article for the 6th edition of the ECSP. Also huge hugs to first time authors from de Piratenpartij, de Volksbank, Leiden University, University of Illinois, Hack in the Box and QuSoft. If IoT was the buzzword in 2017 then Artificial Intelligence (AI) was most definitely in 2018. AI and security seem to be intertwined and that is why you will find several articles about AI in this issue. This year the organization of Hack in the Box created a challenge which you can find at the bottom of the centerfold. There are great prizes involved so make sure to try your luck.
Model-driven engineering is used in the design of systems to (a.o.) enable analysis early in the design process.For instance, by using domain-specific languages, enabling engineers to model systems in terms of their domain, rather then encoding them into general purpose modeling languages.Domain-specific languages, like classical software, evolve over time.When domain languages evolve, they may trigger co-evolution of models, model-to-model transformations, editors (both graphical and textual), and other artifacts that depend on the domain-specific language.This co-evolution can be tedious and very costly.In literature, various approaches are proposed towards automated co-evolution.However, these approaches do not reach full automation.Several other studies have shown that there are theoretical limitations to the level of automation that can be achieved in certain scenarios.For several scenarios full automation can never be achieved.We wish to gain insight to which extent practically occurring scenarios can be automated.To gain this insight, in this paper, we investigate on a large-scale industrial repository, which (co-)evolutionary scenarios occur in practice, and compare them with the various scenarios and their theoretical automatability.We then assess whether practically occurring scenarios can be fully automated.
• A submitted manuscript is the version of the article upon submission and before peer-review. There can be important differences between the submitted version and the official published version of record. People interested in the research are advised to contact the author for the final version of the publication, or visit the DOI to the publisher's website. • The final author version and the galley proof are versions of the publication after peer review. • The final published version features the final layout of the paper including the volume, issue and page numbers.
Forensic analysis of malware activity in network environments is a necessary yet very costly and time consuming part of incident response. Vast amounts of data need to be screened, in a very labor-intensive process, looking for signs indicating how the malware at hand behaves inside e.g., a corporate network. We believe that data reduction and visualization techniques can assist security analysts in studying behavioral patterns in network traffic samples (e.g., PCAP). We argue that the discovery of patterns in this traffic can help us to quickly understand how intrusive behavior such as malware activity unfolds and distinguishes itself from the rest of the traffic.In this paper we present a case study of the visual analytics tool EventPad and illustrate how it is used to gain quick insights in the analysis of PCAP traffic using rules, aggregations, and selections. We show the effectiveness of the tool on real-world data sets involving office traffic and ransomware activity.
System logs typically contain lines with time stamps that each describes an event. Where these events semantically form start and end events, they can be combined into interval events. For visual event analytics, the analysis of interval events is more complex than that of point events, since not only the order of events, but also temporal overlaps have to be taken into account. To address this increased complexity and for the purpose of system understanding and analysis, we present SELE, a domain-independent tool for visualizing parallel interval events. SELE is intended to be used on a single long trace of events. A visual technique named strata timeline is developed to handle visual scalability issues. Finally, a multi-core parallel graph searching algorithm is analyzed to demonstrate SELE.
Multivariate event sequences are ubiquitous: travel history, telecommunication conversations, and server logs are some examples. Besides standard properties such as type and timestamp, events often have other associated multivariate data. Current exploration and analysis methods either focus on the temporal analysis of a single attribute or the structural analysis of the multivariate data only. We present an approach where users can explore event sequences at multivariate and sequential level simultaneously by interactively defining a set of rewrite rules using multivariate regular expressions. Users can store resulting patterns as new types of events or attributes to interactively enrich or simplify event sequences for further investigation. In Eventpad we provide a bottom-up glyph-oriented approach for multivariate event sequence analysis by searching, clustering, and aligning them according to newly defined domain specific properties. We illustrate the effectiveness of our approach with real-world data sets including telecommunication traffic and hospital treatments.
For the protection of critical infrastructures against complex virus attacks, automated network traffic analysis and deep packet inspection are unavoidable. Even with the use of network intrusion detection systems, the number of generated alerts is still too large to analyze manually. In addition, the discovery of domain-specific multi stage viruses (e.g., Advanced Persistent Threats) is typically not captured by a single alert. The result is that security experts are overloaded with low-level technical alerts where they must look for evidence that supports the presence of an APT. In this paper we propose an alert-oriented visual analytics approach for the exploration and analysis of network traffic content. In our approach CoNTA (Contextual analysis of Network Traffic Alerts), experts are supported to discover threats in large alert collections through interactive exploration using selections and attributes of interest. Finally, we show the effectiveness of the approach by applying the approach to real world and artificial data sets.
In this paper we demonstrate how we can study multivariate event sequences in the VAST Mini Challenge 1 data set using our system Eventpad, a notepad editor for event data. We illustrate the effectiveness of multivariate regular expressions, pattern aggregations, and selections to define custom events of interest, discover patterns within sequences, and study differences between sequences. Finally, we discuss our analysis process and summarize some patterns and anomalies we discovered in the data set.
For the protection of critical infrastructures against complex virus attacks, automated network traffic analysis and deep packet inspection are unavoidable. However, even with the use of network intrusion detection systems, the number of alerts is still too large to analyze manually. In addition, the discovery of domain-specific multi stage viruses (e.g., Advanced Persistent Threats) are typically not captured by a single alert. The result is that security experts are overloaded with low-level technical alerts where they must look for the presence of an APT. In this paper we propose an alert-oriented visual analytics approach for the exploration of network traffic content in multiple contexts. In our approach CoNTA (Contextual analysis of Network Traffic Alerts), experts are supported to discover threats in large alert collections through interactive exploration using selections and attributes of interest. Tight integration between machine learning and visualization enables experts to quickly drill down into the alert collection and report false alerts back to the intrusion detection system. Finally, we show the effectiveness of the approach by applying it on real world and artificial data sets.
Most network traffic analysis applications are designed to discover malicious activity by only relying on high-level flow-based message properties. However, to detect security breaches that are specifically designed to target one network (e.g., Advanced Persistent Threats), deep packet inspection and anomaly detection are indispensible. In this paper, we focus on how we can support experts in discovering whether anomalies at message level imply a security risk at network level. In SNAPS (Semantic Network traffic Analysis through Projection and Selection), we provide a bottom-up pixel-oriented approach for network traffic analysis where the expert starts with low-level anomalies and iteratively gains insight in higher level events through the creation of multiple selections of interest in parallel. The tight integration between visualization and machine learning enables the expert to iteratively refine anomaly scores, making the approach suitable for both post-traffic analysis and online monitoring tasks. To illustrate the effectiveness of this approach, we present example explorations on two real-world data sets for the detection and understanding of potential Advanced Persistent Threats in progress.
• A submitted manuscript is the author's version of the article upon submission and before peer-review. There can be important differences between the submitted version and the official published version of record. People interested in the research are advised to contact the author for the final version of the publication, or visit the DOI to the publisher's website. • The final author version and the galley proof are versions of the publication after peer review. • The final published version features the final layout of the paper including the volume, issue and page numbers.
Mark G. J. Van Den Brand合作论文数Eindhoven University of Technology
Mathematics and Computer Science1