
This paper describes the use of machine learning techniques to implement a Bayesian approach to modelling the dependency between offence data and environmental factors such as demographic characteristics and spatial location. The main goal of this paper is to provide a fully probabilistic approach to modelling crime which reflects all uncertainties in the prediction of offences as well as the uncertainties surrounding model parameters.
This paper describes an approach for detecting the presence or emergence of organised crime (OC) signals on social media. It shows how words and phrases, used by members of the public in social media posts, can be treated as weak signals of OC, enabling information to be classified according to a taxonomy. Formal concept analysis is used to group information sources, according to crime-type and location, thus providing a means of corroboration and creating OC concepts that can be used to alert police analysts to the possible presence of OC. The analyst is able to ‘drill down’ into an OC concept of interest, discovering additional information that may be pertinent to the crime. The paper describes the implementation of this approach into a fully-functional prototype software system, incorporating a social media scanning system and a map-based user interface. The approach and system are illustrated using human trafficking and modern slavery as an example. Real data is used to obtain results that show that weak signals of OC have been detected and corroborated, thus alerting to the possible presence of OC.
Law enforcement and intelligence agencies generally have access to a number of rich data sources, both structured and unstructured, and with the advent of high performing entity resolution it is now possible to fuse multiple heterogeneous datasets into an explicit generic data representation. But once this is achieved how should agencies go about attempting to exploit this data by proactively identifying criminal events and the actors responsible? The authors will outline an effective generic method that; computationally extracts minimally overlapping contextual subgraphs, then uses these subgraphs as the basis to construct a mesoscopic graph based on the intersections between the subgraphs, enabling knowledge discovery from these data representations for the purpose of maximally disrupting terrorism, organised crime and the broader criminal network.
The information visualization of networks has been a tricky task during the last decade. Visualization of features by way of merging, linking, and grouping of entity attributes is provided to criminal network investigators. It is difficult to understand such large amounts of statistical data. A number of solutions have been proposed to tackle this bulk of information. We have found that the prevailing challenges to information visualization can be eliminated to a large extent by detecting evolving network patterns which are extracted by way of visual analysis of criminal activity based on temporal data, by examining some dynamics of criminal networks, and by making use of some novel interactive features. The current study will help to understand interesting patterns in criminal data by way of visualization. Besides our previously proposed network visualization features, we have appended five new features. These features include ‘Pie-chart feature’ which has been proposed for better ‘details on demand’ facility to the analysts. A ‘Trend analysis feature’ is proposed for visualizing the variation in different crimes over some span of time. The ‘Graphical Trend Analysis feature’ provides a graphical interface to the analysts. There is a unique ‘Encircle feature’, with the aid of which the desired clusters can be dragged away from the dense network for easy manipulation. With ‘Similar node feature’, the analysts may get summarized information regarding the activity of different nodes which are at distant apart. We have made an evaluation of our proposed visualization features by conducting an experiment. Thirty-two participants evaluated the system. The experiment was performed in two phases. In the first phase, a usability evaluation and qualitative feedback was carried out to check whether the features provided adequate results to the users. In the second phase, the comparison of the features had been performed against some other state-of-the-art tool. These tasks were to be performed in the groups of participants. The public data set of Chicago Narcotics was used. We found that the participants, of the PEVNET group, performed the tasks faster as compared to the other techniques used in the experiment. We have demonstrated the usability of the new features with examples by employing the datasets. We have proposed a unique way of visualizing the clustering of data, with which the analyst gets a sound visualization of the data. The usability, of the proposed features, indicates that the crime analysts will get a valuable insight into the criminal networks.
Human trafficking is one of the most atrocious crimes and among the challenging problems facing law enforcement which demands attention of global magnitude. In this study, we leverage textual data from the website “Backpage”—used for classified advertisement—to discern potential patterns of human trafficking activities which manifest online and identify advertisements of high interest to law enforcement. Due to the lack of ground truth, we rely on a human analyst from law enforcement, for hand-labeling a small portion of the crawled data. We extend the existing Laplacian SVM and present S^3VM-R , by adding a regularization term to exploit exogenous information embedded in our feature space in favor of the task at hand. We train the proposed method using labeled and unlabeled data and evaluate it on a fraction of the unlabeled data, herein referred to as unseen data, with our expert’s further verification. Results from comparisons between our method and other semi-supervised and supervised approaches on the labeled data demonstrate that our learner is effective in identifying advertisements of high interest to law enforcement.
Attackers increasingly take advantage of naive users who tend to treat non-executable files casually, as if they are benign. Such users often open non-executable files although they can conceal and perform malicious operations. Existing defensive solutions currently used by organizations prevent executable files from entering organizational networks via web browsers or email messages. Therefore, recent advanced persistent threat attacks tend to leverage non-executable files such as portable document format (PDF) documents which are used daily by organizations. Machine Learning (ML) methods have recently been applied to detect malicious PDF files, however these techniques lack an essential element—they cannot be efficiently updated daily. In this study we present an active learning (AL) based framework, specifically designed to efficiently assist anti-virus vendors focus their analytical efforts aimed at acquiring novel malicious content. This focus is accomplished by identifying and acquiring both new PDF files that are most likely malicious and informative benign PDF documents. These files are used for retraining and enhancing the knowledge stores of both the detection model and anti-virus. We propose two AL based methods: exploitation and combination. Our methods are evaluated and compared to existing AL method (SVM-margin) and to random sampling for 10 days, and results indicate that on the last day of the experiment, combination outperformed all of the other methods, enriching the signature repository of the anti-virus with almost seven times more new malicious PDF files, while each day improving the detection model’s capabilities further. At the same time, it dramatically reduces security experts’ efforts by 75 %. Despite this significant reduction, results also indicate that our framework better detects new malicious PDF files than leading anti-virus tools commonly used by organizations for protection against malicious PDF files.
The research presented, investigates the optimal set of operational codes (opcodes) that create a robust indicator of malicious software (malware) and also determines a program’s execution duration for accurate classification of benign and malicious software. The features extracted from the dataset are opcode density histograms, extracted during the program execution. The classifier used is a support vector machine and is configured to select those features to produce the optimal classification of malware over different program run lengths. The findings demonstrate that malware can be detected using dynamic analysis with relatively few opcodes.
Currently 40 % of the world’s population, around 3 billion users, are online using cyberspace for everything from work to pleasure. While there are numerous benefits accompanying this medium, the Internet is not without its perils. In this case study article, we focus specifically on the challenge of fake (or unnatural) online identities, such as those used to defraud people and organisations, with the aim of exploring an approach to detect them.
Social media channels, such as Facebook or Twitter, allow for people to express their views and opinions about any public topics. Public sentiment related to future events, such as demonstrations or parades, indicate public attitude and therefore may be applied while trying to estimate the level of disruption and disorder during such events. Consequently, sentiment analysis of social media content may be of interest for different organisations, especially in security and law enforcement sectors. This paper presents a new lexicon-based sentiment analysis algorithm that has been designed with the main focus on real time Twitter content analysis. The algorithm consists of two key components, namely sentiment normalisation and evidence-based combination function, which have been used in order to estimate the intensity of the sentiment rather than positive/negative label and to support the mixed sentiment classification process. Finally, we illustrate a case study examining the relation between negative sentiment of twitter posts related to English Defence League and the level of disorder during the organisation’s related events.
Effectively assessing and configuring security controls to minimize network risks requires human judgment. Little is known about what factors network professionals perceive to make judgments of network risk. The purpose of this research was to examine first, what factors are important to network risk judgments (Study 1) and second, how risky/safe each factor is judged (Study 2) by a sample of network professionals. In Study 1, a complete list of factors was generated using a focus group method and validated on a broader sample using a survey method with network professionals. Factors detailing the adversary and organizational network readiness were rated highly important. Study 2 investigated the level of riskiness for each factor that is described in a vignette-based factor scenario. The vignette provided context that was missing in Study 1. The highest riskiness ratings were of factors detailing the adversary and the lowest riskiness ratings detailed the organizational network readiness. A significant relationships existed in Study 2 between the level of agreement on each factor's rating across our sample of network professionals and the riskiness level each factor was judged. Factors detailing the adversary were highly agreed upon while factors detailing the organizational capability were less agreed upon. Computational risk models and network risk metrics ask professionals to perceive factors and judge overall network risk levels but no published research exists on what factors are important for network risk judgments. These empirical findings address this gap and factors used in models and metrics could be compared to factors generated herein. Future research and implications are discussed at the close of this paper.
In order to provide a foundation for education on e-discovery and security in Electronic Health Record (EHR) systems, this paper identifies emerging issues in the area. Based on a detailed literature review it details key categories: Development in EHR, E-discovery policy and strategy, and Security and privacy in EHR and also discusses e-discovery issues in cloud computing and big data contexts. This may help to create a framework for potential short course-design on e-discovery and security in the healthcare domain.
DNSSEC offers protection against spoofing of DNS data by providing origin authentication, ensuring data integrity and authentication of non-existence by using public-key cryptography. Although the relevance of securing a technology as crucial to the Internet as DNS is obvious, the DNSSEC implementation increases the complexity of the deployed DNS infrastructure, which may result in misconfiguration. In this article, we measure and analyze the misconfigurations for domains in six zones (.bg, .br, .co, .com, .nl and .se). Furthermore, we categorize these misconfigurations and provide an explanation for their possible causes. Finally, we evaluate the effects of misconfigurations on the reachability of a zone’s network. Our results show that, although progress has been made in the implementation of DNSSEC, over 4 % of evaluated domains show misconfigurations. The domains with the most frequently appearing misconfiguration are often hosted at a very limited set of hosting providers. Of these misconfigured domains, almost 75 % were unreachable from a DNSSEC-aware resolver. This illustrates that although the authorities of a domain may think their DNS is secured, it is in fact not. Worse still, misconfigured domains are at risk of being unreachable from the clients who care about and implement DNSSEC verification, while the publisher may remain unaware of the error and its consequences.
Wildlife trafficking, a focus of organized transnational crime syndicates, is a threat to biodiversity. Such crime networks span beyond protected areas holding strongholds of species of interest such as African rhinos. Such networks extend over several countries and hence beyond the jurisdiction of any one law enforcement authority. We show how a federated database can overcome disjoint information kept in different databases. We also show how social network analyses can provide law enforcers with targeted responses that maximally disrupt a criminal network. We introduce an actionable intelligence report using social network measures that identifies key players and predicts player succession. Using a rhino case study we illustrate how such a report can be used to optimize enforcement operations.
Many people who discuss sensitive or private issues on social media services are using pseudonyms or aliases in order to not reveal their true identity, while using their usual, non-private accounts when posting messages on less sensitive issues. Previous research has shown that if those individuals post large amounts of user-generated content, stylometric techniques can be used to identify the author based on the characteristics of the textual content. In this article we show how an author’s identity can be unmasked in a similar way using various time features (e.g., period of the day and the day of the week when a user’s posts have been published). We combine several different time features into a timeprint , which can be seen as a type of fingerprint when identifying users on social media. We use supervised machine learning (i.e., author identification) and unsupervised alias matching (similarity detection) in a number of different experiments with forum data to get an understanding of to what extent timeprints can be used for identifying users in social media, both in isolation and when combined with stylometric features. The obtained results show that timeprints indeed can be a very powerful tool for both author identification and alias matching in social media.
Duplicate and false identity records are quite common in identity management systems due to unintentional errors or intentional deceptions. Identity resolution is to uncover identity records that are co-referent to the same real-world individual. In this paper we introduce a framework of identity resolution that covers different identity attributes and matching algorithms. Guided by social identity theories, we define three types of identity cues, namely personal identity attributes, social behavior attributes, and social relationship attributes. We also compare three matching algorithms: pair-wise comparison, transitive closure, and collective clustering. Our experiments using synthetic and real-world data demonstrate the importance of social behavior and relationship attributes for identity resolution. In particular, a collective identity resolution technique, which captures all three types of identity attributes and makes matching decisions on identities collectively, is shown to achieve the best performance among all approaches.
Popular network scan detection algorithms operate through evaluating external sources for unusual connection patterns and traffic rates. Research has revealed evasive tactics that enable full circumvention of existing approaches (specifically the widely cited Threshold Random Walk algorithm). To prevent use of these circumvention techniques, we propose a novel approach to network scan detection that evaluates the behavior of internal network nodes, and combine it with other established techniques of scan detection. By itself, our algorithm is an efficient, protocol-agnostic, completely unsupervised method that requires no a priori knowledge of the network being defended beyond which hosts are internal and which hosts are external to the network, and is capable of detecting network scanning attempts regardless of the rate of the scan (working even with connectionless protocols). We demonstrate the effectiveness of our method on both live data from an enterprise-scale network and on simulated scan data, finding a false positive rate of just 0.000034% with respect to the number of inbound flows. When combined with both Threshold Random Walk and simple rate-limiting detection, we achieve an overall detection rate of 94.44%.
Editorial Criminal networks such as terrorist and organized crime networks pose a threat to both national and international security. Investigation of such networks by law enforcement and intelligence agencies is a complex and time-consuming task. The ability to understand and model such networks, to analyze and visualize such networks, to mine and forecast such networks, to predict and simulate such networks, and to effectively collaborate and share information about such networks are key factors in defeating them. It is well known that the Internet is used for criminal activities such as terrorism. The Center for Terror Analysis, Danish Security and Intelligence Service wrote a report in 2008 that states: The Internet plays a significant and increasing role for militant extremists and terrorist groups that use this media to spread messages, communicate and carry out virtual training as well as for recruitment, logistic support and operational preparations [1]. The United Nations Office on Drugs and Crime (UNODC) wrote a report in 2012 that states: Technology is one of the strategic factors driving the increasing use of the Internet by terrorist organizations and their supporters for a wide range of purposes, including recruitment, financing, propaganda, training, incitement to commit acts of terrorism, and the gathering and dissemination of information for terrorist purposes [2]. Two papers of this special issue focus on matters related to the use of the Internet for criminal activities. The first paper in this special issue by Torok (Developing an explanatory model for the process of online radicalization and terrorism) proposes a model for understanding the mechanisms and power relations that underlie the use of Internet services such as Facebook, YouTube, and Twitter for online radicalization and terrorism. The model is based on previous work by Foucault on psychiatric power [3]. The second paper in this special issue by BowmanGrieve (A psychological perspective on virtual communities supporting terrorist & extremist ideologies as a tool for recruitment) examines the role of virtual communities as a tool for recruitment used by terrorist and extremist organizations. A case study looks at the how the Radical Right uses the Internet in their activities. Many terrorist organizations are transnational in nature [2]. Hence, it is important to understand and predict the locational dynamics of such organizations. The third paper in this special issue by Desmarais and Cranmer (Forecasting the locational dynamics of transnational terrorism: a network analytic approach) proposes and evaluates a network analytic approach to predict the geopolitical sources and targets of terrorism. Software tools play an increasing role in criminal network investigation [4]. Software tools can help support and (semi) automate some of the complex knowledge management processes involved in criminal network investigation (collection, processing, synthesis, sense-making, dissemination, etc.). The fourth and final paper in this special issue by Petersen and Wiil (CrimeFighter Investigator: Integrating synthesis and sense-making for criminal network investigation) presents a new software tool dedicated to support criminal network investigation.
This paper adopts the metaphor of representational fluency and proposes an auto linking approach to help analysts investigate details of suspicious sections across different cybersecurity visualizations. Analysis of spatiotemporal network security data takes place both conditionally and in sequence. Many visual analytics systems use time series curves to visualize the data from the temporal perspective and maps to show the spatial information. To identify anomalies, the analysts frequently shift across different visualizations and the original data view. We consider them as various representations of the same data and aim to enhance the fluency of navigation across these representations. With the auto linking mechanism, after the analyst selects a segment of a curve, the system can automatically highlight the related area on the map for further investigation, and the selections on the map or the data views can also trigger the related time series curves. This approach adopts the slicing operation of the Online Analytical Process (OLAP) to find the basic granularities that contribute to the overall value change. We implemented this approach in an award-winning visual analytics system, SemanticPrism, and demonstrate the functions through two use cases.
Insurgency emerges from many interactions between numerous social, economical, and geographical factors. Adequately accounting for the large number of potentially relevant interactions, and the complex ways in which they operate, is key to creating valuable models of insurgency. However, this has long been a challenging endeavour, as insurgency imposes specific limitations on the data that could speak to these interactions: quantitative data is limited by the difficulties of systematic collection in war, while qualitative data may include vague or conflicting insights from direct observers. In this paper, we designed a computational framework based on Fuzzy Cognitive Maps and Complex Networks to face these limitations. A software solution fully implements this framework and allows analysts to conduct simulations, in order to better understand the current dynamics of insurgency or test ‘what-if’ scenarios. Two approaches are presented to guide analysts in developing models based on our framework, either through a nuanced reading of the literature, or by aggregating the knowledge of domain experts.