In absence of sufficient data, structured expert judgment is a suitable method to estimate uncertain quantities. While such methods are well established for individual variables, eliciting their dependence in a structured manner is a less explored field of research. We tested the performance of experts in constructing and quantifying a nonparametric Bayesian network, describing the correlation between river tributary discharges. Specialized software was provided to assist the experts. Expert performance was investigated using the dependence calibration score (a correlation matrix distance metric) and the likelihood of the joint distribution. Desirable properties of the dependence calibration score were investigated theoretically. Individual expert judgments were combined based on performance into a group opinion aka decision maker. All experts were able to create and quantify a correlation matrix between 10 variables that resembled the correlations between observed discharges well. The decision makers performed similarly to the best expert. Based on the metrics investigated, it mattered little which expert opinions and with what weight were combined in a decision maker. This is partly because all experts performed well. Adding a bad performing expert increased the positive effect of performance-based weighting, underscoring the importance of developing scoring rules for dependence elicitation. The overall results are promising: Aided by specialized graphical software, the experts in this study were able to quickly create and quantify dependence structures.
In this paper we discuss PyBanshee, which is a Python-based open-source implementation of the MATLAB toolbox BANSHEE. PyBanshee constitutes the first fully open-source package to quantify, visualize and validate Non-Parametric Bayesian Networks (NPBNs). The architecture of PyBanshee is heavily based on its MATLAB predecessor. It presents the full implementation of existing tools and introduces new modules. Specifically, PyBanshee allows for: (i) choosing fully parametric one-dimensional margins, (ii) choosing different sample sizes for the model-validation tests based on the Hellinger distance, (iii) drawing user-defined sample sizes of the NPBN, (iv) sample-based conditioning sampling (similarly to the closed-source proprietary package UNINET by LightTwist Software) and (v) visualizing the comparison between the histograms of the unconditional and conditional marginal distributions. New detailed examples demonstrating new features are provided.
Background: Clinical decision support systems (CDSS) are a category of health information technologies that can assist clinicians to choose optimal treatments. These support systems are based on clinical trials and expert knowledge; however, the amount of data available to these systems is limited. For this reason, CDSSs could be significantly improved by using the knowledge obtained by treating patients. This knowledge is mainly contained in patient records, whose usage is restricted due to privacy and confidentiality constraints. Methods: A treatment effectiveness measure, containing valuable information for treatment prescription, was defined and a method to extract this measure from patient records was developed. This method uses an advanced cryptographic technology, known as secure Multiparty Computation (henceforth referred to as MPC), to preserve the privacy of the patient records and the confidentiality of the clinicians' decisions. Results: Our solution enables to compute the effectiveness measure of a treatment based on patient records, while preserving privacy. Moreover, clinicians are not burdened with the computational and communication costs introduced by the privacy-preserving techniques that are used. Our system is able to compute the effectiveness of 100 treatments for a specific patient in less than 24 minutes, querying a database containing 20,000 patient records. Conclusion: This paper presents a novel and efficient clinical decision support system, that harnesses the potential and insights acquired from treatment data, while preserving the privacy of patient records and the confidentiality of clinician decisions.
Bayesian Networks (BNs) are probabilistic, graphical models for representing complex dependency structures. They have many applications in science and engineering. Their particularly powerful variant – Non-Parametric BNs – are for the first time implemented as an open-access scriptable code, in the form of a MATLAB toolbox “BANSHEE”.1 1 BANSHEE stands for ‘Bayesian Networks in Scholarly Endeavours’. However, a banshee is also, in Irish folklore, a female spirit whose appearance is a warning about impending death. Bayesian Networks have been extensively used in risk analysis in fields ranging from aviation safety through natural hazards to building fire safety, hence they also warn against possible dangers, somewhat similarly to a banshee. The software allows for quantifying the BN, validating the underlying assumptions of the model, visualizing the network and its corresponding rank correlation matrix, and finally making inference with a BN based on existing or new evidence. We also include in the toolbox, and discuss in the paper, some applied BN models published in most recent scientific literature.
Various equicontinuity properties for families of Markov operators have been – and still are – used in the study of existence and uniqueness of invariant probability for these operators, and of asymptotic stability. We prove a general result on equivalence of equicontinuity concepts. It allows comparing results in the literature and switching from one view on equicontinuity to another, which is technically convenient in proofs. More precisely, the characterisation is based on a ‘Schur-like property’ for measures: if a sequence of finite signed Borel measures on a Polish space is such that it is bounded in total variation norm and such that for each bounded Lipschitz function the sequence of integrals of this function with respect to these measures converges, then the sequence converges in dual bounded Lipschitz norm to a measure.
Collaboration between financial institutions helps to improve detection of fraud. However, exchange of relevant data between these institutions is often not possible due to privacy constraints and data confidentiality. An important example of relevant data for fraud detection is given by a transaction graph, where the nodes represent bank accounts and the links consist of the transactions between these accounts. Previous works show that features derived from such graphs, like PageRank, can be used to improve fraud detection. However, each institution can only see a part of the whole transaction graph, corresponding to the accounts of its own customers. In this research a new method is described, making use of secure multiparty computation (MPC) techniques, allowing multiple parties to jointly compute the PageRank values of their combined transaction graphs securely, while guaranteeing that each party only learns the PageRank values of its own accounts and nothing about the other transaction graphs. In our experiments this method is applied to graphs containing up to tens of thousands of nodes. The execution time scales linearly with the number of nodes, and the method is highly parallelizable. Secure multiparty PageRank is feasible in a realistic setting with millions of nodes per party by extrapolating the results from our experiments.
Water distribution networks (WDNs) are critical to provide safe, clean drinking water around the globe. However, they are susceptible to accidental or deliberate contamination, potentially resulting in poisoned water, many fatalities and large economic consequences. In order to protect against such intrusions, an efficient sensor network should be placed in a WDN. Finding the optimal placement for water quality sensors is a challenging problem. Several sensor placement strategies have been proposed, but the vast majority of these strategies rely on the assumption that the sensors are perfect. In this paper we provide evidence for the imperfection of water quality sensors, by conducting measurements in an operational environment. We investigate the imperfection of four types of water quality sensors being employed in actual WDNs for the purpose of contamination detection. We describe experiments conducted at the WaDi testbed, a realistic water distribution facility at the Singapore University of Technology and Design. Through these experiments we study the imperfection, sensitivity and degradation of the water quality sensors, under normal conditions (water flow without contaminants present) as well as under attack conditions. It is shown that several aspects of sensor imperfection do occur, including missing values, inexplicable jumps and drifting.
Water Distribution Networks (WDNs) are often susceptible to either accidental or deliberate contamination which can lead to poisoned water, many fatalities and large economic consequences. In order to protect against these intrusions or attacks, an efficient sensor network with a limited number of sensors should be placed in a WDN. In this paper, we focus on optimal sensor placements by introducing two greedy-based algorithms in which the imperfection of sensors and multiple objectives can be taken into account. The algorithms were tested using a medium scale urban WDN. It is shown that our algorithms are able to find sensor placements in reasonable time and that its solutions are close to optimal. Furthermore, relaxing the often used assumption that sensors work perfectly results in different sensor placements than were found before, indicating the importance to take sensor imperfection into account when placing sensors. (C) 2018 Elsevier Ltd. All rights reserved.
A Banach space has the Schur property when every weakly convergent sequence converges in norm. We prove a Schur-like property for measures: if a sequence of finite signed Borel measures on a Polish space is such that it is bounded in total variation norm and such that for each bounded Lipschitz function the sequence of integrals of this function with respect to these measures converges, then the sequence converges in dual bounded Lipschitz norm or Fortet-Mourier norm to a measure. Moreover, we prove three consequences of this result: the first is equivalence of concepts of equicontinuity in the theory of Markov operators, the second is the derivation of weak sequential completeness of the space of signed Borel measures on Polish spaces from our main result and the third concerns conditions for the coincidence of weak and norm topologies on sets of measures that are bounded in total variation norm with additional properties.
In this paper two methodologies are investigated that contribute to better assessment of risks related to extreme rainfall events. Firstly, one-parameter bivariate copulas are used to analyze rain gauge data in the Netherlands. Out of three models considered, the Gumbel copula, which indicates upper tail dependence, represents the data most accurately for all 33 stations in the Netherlands. Seasonal variability is noticeable, with rank correlation reaching maximum in winter and minimum in summer as well as other temporal and spatial patterns. Secondly, an expert judgment elicitation was undertaken. The experts' opinions were combined using Cooke's classical method in order to obtain estimates of future changes in precipitation patterns. Experts predicted mostly an approximate 10% increase in rain amount, duration, intensity and the dependence between amount and duration. The results were in line with official national climate change scenarios, based on numerical modelling. Applicability of both methods was presented based on an example of an existing tunnel in the Netherlands, contributing to better estimates of the tunnel's limit state function and therefore the probability of failure. (c) 2017 American Society of Civil Engineers.
Geometric models are a proven aid for calculating the CapEx needed for a network deployment and used widely by telecommunication network operators for their access networks. However, for the economic viability an accurate estimation of the delivered bandwidth is important as well. In this paper a method is described to calculate a bandwidth profile for an area under investigation, given a single topology and using the geometric representation. Next to this single topology result, a method to estimate the bandwidth profile for a geometric model of a multi-layer network topology is presented.
In this paper we take one of the cutting edge algorithms for computing the all-terminal reliability and the k-terminal reliability of a network and use it to compute the reliability of a real life gas distribution network in the Netherlands. To do this we estimate network properties using industry knowledge and combine several different techniques to make the problem computable. This is the first time known to us that these techniques have been applied to a large, in this case over 20000 nodes, real life network. Besides this, we show the versatility of this pathwidth-based dynamic programming algorithm by suggesting some powerful but simple modifications and argue that this network is representative for other distribution networks.
The online environment offers a fertile breeding ground for anti-brand herds of disgruntled consumers. Firms are often caught off guard by the unpredictability of such herds and, as a consequence, are forced into a reactive, defensive stance. We conduct a social media analysis that aims to shed light on the formation, growth, and dissolution of online anti-brand herds. First we expand on the concept of environmental turbulence to advance core properties unique to online herd behavior. Next, based on evidence gathered from 40 online anti-brand herd episodes targeting two prominent firms from the Netherlands, we develop an analytical model to investigate drivers of herd formation, growth, and dissolution. Finally, combining environmental turbulence literature with our empirical findings, we derive a novel typology of online anti-brand herd behaviors, and put forward six propositions to guide theory development in this area.
This paper deals with the problem of inferring short time-scale fluctuations of a system's behavior from periodic state measurements. In particular, we devise a novel, efficient procedure to compute four interesting performance metrics for a transient birth-death process on an interval of fixed length with given begin and end states: the probability to exceed a predefined (critical) level in, and the expectation of the time, area, and number of arrivals above level m. Moreover, our procedure allows to compute the variances and cross-correlations of the latter three metrics. The asymptotic behavior of the metrics for small and large measurement intervals is also derived.An extensive numerical study in the context of communication networks reveals the impact of important system parameters on the considered performance metrics, and shows that the three latter metrics are very highly correlated. We also illustrate through this numerical study how our analysis can be used in practical situations to support e.g., capacity management and SLA verification. (C) 2015 Elsevier B.V. All rights reserved.
The aim of this paper is to improve scientific modeling of interdependent socio-technical networks. In these networks the interplay between technical or infrastructural elements on the one hand and social and behavioral aspects on the other hand, plays an important role. Examples include electricity networks, financial networks, residential choice networks. We propose an Agent-Based Modeling approach to simulate interdependent technical and social network behavior, the effects of potential policy measures and the societal impact when disturbances occur, where we focus on a specific use case: the smart grid, an intelligent system for matching supply and demand of electricity.
In this note, we consider Feller transition functions (P-t)(t is an element of[0,+infinity)) defined on a Polish space (X,d), the associated families ((S-t,T-t))(t is an element of[0),(+infinity)) of Markov-Feller pairs, we think of (S-t)(t is an element of[0,+infinity)) as a semigroup of positive contractions of C-b(X) = the Banach space of all real-valued bounded continuous functions defined on X, and we assume that (S-t)(t is an element of[0,+infinity)) has a generator A defined on the entire C-b(X). For this type of transition function, we characterize completely the sets and that appear in the KBBY (Krylov-Bogolioubov-Beboutoff-Yosida) ergodic decomposition defined by (P-t)(t is an element of[0,+infinity)) in terms of A, only. The above-mentioned characterizations of and are a first step in a research program that consists of articulating the KBBY decomposition in terms of the generator and then using the results in order to study the decomposition for various continuous-time time-homogeneous Markov processes (for a description of this research direction, see the subsection 2. Transition functions of Markov processes in section 4. Future research of the second author's paper Transition probabilities, transition functions, and an ergodic decomposition, Bull. of the Transilvania Univ. of Braffsov, Vol. 15(50), Series B-2008; see also Introduction of the second author's monograph Invariant Probabilities for Transition Functions, Springer, 2014. We then use the characterizations of the sets Gamma(epi) and Gamma(epie) in terms of A in order to study certain exponential one-parameter convolution semigroups of probability measures, and to extend and strengthen a result of Hunt discussed in Heyer's 1977 monograph Probability Measures on Locally Compact Groups.
We propose graph theoretical resilience metrics for the identification of critical pipes in gas distribution networks, i.e. pipes whose failure would have the greatest impact in terms of gas demand not delivered. These metrics are applied to a case study and compared to realistic gas flow simulations that incorporate pipe failures. We succeed in identifying the critical pipes by using a metric closely related to efficiency.
Many threats in the real world can be related to activities in public sources on the Internet. Early detection of threats based on Internet information could assist in the prevention of incidents. However, the amount of data in social media, blogs and forums rapidly increases and it is time consuming for security services to monitor all these sources. Therefore, it is important to have a system that automatically ranks messages based on their threat potential and thereby allows security operators to check these messages more efficiently. In this paper, we present a novel method for detecting threatening messages on Twitter based on trigger keywords and contextual cues. The system was tested on multiple large collections of Dutch tweets. Our experimental results show that our system can successfully analyze messages and recognize threatening content.
N. Litvak合作论文数Faculty of Electrical Engineering, Mathematics and Computer Science
University of Twente1