Assessing the disclosure risk of releasing the result of a data processing activity is inherently difficult: a data processing activity might consist of several steps with complicated interleavings; it is not clear on which personal properties to concentrate and what the severity implication is of disclosing each of these properties to each affected individual; it is not clear which kind of additional information is or will be available to an attacker and how the results are further processed. This paper proposes a disclosure risk assessment strategy that tackles these challenges. The proposed risk assessment strategy is mathematically rigorous, enables modular assessment and only needs to concentrate on the membership property. The latter expresses whether an attacker can reason whether an individual’s data has been used in the data processing activity. This risk provides as an upper bound for the severity for various personal properties for different individuals such that there is no need to enumerate and discuss them separately. The resulting disclosure risk assessment holds even if additional information is available to an attacker and even for the result of a further processing step based solely on public or previously released data, and can be conveniently applied to various threat assumptions. Our risk assessment notion takes the perspective of estimating how well a person is hidden within a crowd after the result of a data processing activity is released. Such crowd sizes have turned out to be more comprehensible than other, more probabilistic, notions. At its core, this paper presents a mathematically rigorous translation from Differential Privacy to a comprehensive and comprehensible disclosure risk assessment notion (for technical means): the ratio (preservation factor) between the size of the imaginary crowd in which any person whose data is used is hidden before and after observing a data release. We show that releasing the result of a data processing activity that solely executes an epsilon-DP algorithm results in a crowd size preservation factor of e^-ε .
Mixnets are widely believed to hide communication metadata of individuals. We show that there are various pitfalls when designing mixnet topologies and routing strategies, in particular when choosing mixnets with low delays. We introduce a tool that empirically evaluates such leakage in mixnets and show that this tool precisely estimates this leakage for recipient anonymity, up to an error introduced by sampling. First, we introduce a novel generic attack strategy that we even prove to be optimal for breaking recipient anonymity. In contrast to prior work, our attack strategy incorporates the severity of each observation's leakage, via its so-called privacy loss. Second, our tool provides a lower bound on an attacker's advantage against recipient anonymity by sampling a large set of observations; if a significant number of observations with high privacy loss is observed, the tool outputs a lower bound on the leakage by providing a lower bound on the mass of the tail of the distribution of privacy losses. From the literature, we study the topology and routing strategies of the Karaoke and Atom protocols, provide bounds on their leakage, and recommend design choices based on the analysis.
Differentially private massively distributed learning poses one key challenge when compared to differentially private centralized learning, where all data are aggregated at one party: minimizing communication overhead while achieving strong utility-privacy tradeoffs. The minimal amount of communication for distributed learning is non-interactive communication, i.e., each party only sends one message. In this work, we propose two differentially private, non-interactive, distributed learning algorithms in a framework called Secure Distributed \helmet. This framework is based on what we coin blind averaging: each party locally learns and noises a model and all parties then jointly compute the mean of their models via a secure summation protocol (e.g., secure multiparty computation). The learning algorithms we consider for blind averaging are empirical risk minimizers (ERM) like SVMs and Softmax-activated single-layer perception (Softmax-SLP). We show that blind averaging preserves privacy if the models are averaged via secure summation and the objective function is smooth, Lipschitz, and strongly convex. We show that the objective function of Softmax-SLP fulfills these criteria, which implies leave-one-out robustness and might be of independent interest. On the practical side, we provide experimental evidence that blind averaging for SVMs and Softmax-SLP can have a strong utility-privacy tradeoff: we reach an accuracy of $86$ \% on CIFAR-10 for $\varepsilon = 0.36$ and $1{,}000$ users and of $44$ \% on CIFAR-100 for $\varepsilon = 1.18$ and $100$ users, both after a SimCLR-based pre-training. As an ablation, we study the resilience of our approach to a strongly non-IID setting. On the theoretical side, we show that in the limit blind averaging hinge-loss based SVMs convergences to the centralized learned SVM. Our approach is based on the representer theorem and can be seen as a blueprint for finding convergence for other ERM problems like Softmax-SLP.
One of the main goals of financial institutions (FIs) today is combating fraud and financial crime. To this end, FIs use sophisticated machine-learning models trained using data collected from their customers. The output of machine learning models may be manually reviewed for critical use cases, e.g., determining the likelihood of a transaction being anomalous and the subsequent course of action. While advanced machine learning models greatly aid an FI in anomaly detection, model performance could be significantly improved using additional customer data from other FIs. In practice, however, an FI may not have appropriate consent from customers to share their data with other FIs. Additionally, data privacy regulations may prohibit FIs from sharing clients' sensitive data in certain geographies. Combining customer data to jointly train highly accurate anomaly detection models is therefore challenging for FIs in operational settings. In this paper, we describe a privacy-preserving framework that allows FIs to jointly train highly accurate anomaly detection models. The framework combines the concept of federated learning with efficient multi-party computation and noisy aggregates inspired by differential privacy. The presented framework was submitted as a winning entry to the financial crime detection track of the US/UK PETs Challenge. The challenge considered an architecture where banks hold customer data and execute transactions through a central network. We show that our solution enables the network to train a highly accurate anomaly detection model while preserving privacy of customer data. Experimental results demonstrate that use of additional customer data using the proposed approach results in improvement of our anomaly detection model's AUPRC from 0.6 to 0.7. We discuss how our framework, can be generalized to other similar scenarios.
Users embrace the rapid development of virtual reality (VR) technology for everyday settings. These settings include payment, which makes user authentication necessary. Despite this need, there is a limited understanding of how users' unique experiences in VR contribute to their security perception. To understand this question, we designed probes of payment authentication, which are embedded in the routine payment of a VR game, to provoke participants' reactions from multiple angles.
While many anonymous communication (AC) protocols have been proposed to provide anonymity over the internet, scaling to a large number of users while remaining provably secure is challenging. We tackle this challenge by proposing a new scaling technique to improve the scalability/anonymity of AC protocols that distributes the computational load over many nodes without completely disconnecting the paths different messages take through the network. We demonstrate that our scaling technique is useful and practical through a core sample anonymous broadcast protocol, Streams, that offers provable security guarantees and scales for a million messages. The scaling technique ensures that each node in the system does the computation-heavy public key operation only for a tiny fraction of the total messages routed through the Streams network while maximizing the mixing/shuffling in every round. Our experimental results show that Streams can scale well even if the system has a load of one million messages at any point in time, with a latency of 16 seconds while offering provable “one-in-a-billion” unlinkability, and can be leveraged for applications such as anonymous microblogging and network-level anonymity for blockchains. We also illustrate by examples that our scaling technique can be useful to other AC protocols to improve their scalability and privacy, and can be interesting to protocol developers.
Users readily embrace the rapid advancements in virtual reality (VR) technology within various everyday contexts, such as gaming, social interactions, shopping, and commerce. In order to facilitate transactions and payments, VR systems require access to sensitive user data and assets, which consequently necessitates user authentication. However, there exists a limited understanding regarding how users' unique experiences in VR contribute to their perception of security. In our study, we adopt a research approach known as ``technology probe'' to investigate this question. Specifically, we have designed probes that explore the authentication process in VR, aiming to elicit responses from participants from multiple perspectives. These probes were seamlessly integrated into the routine payment system of a VR game, thereby establishing an organic study environment. Through qualitative analysis, we uncover the interplay between participants' interaction experiences and their security perception. Remarkably, despite encountering unique challenges in usability during VR interactions, our participants found the intuitive virtualized authentication process beneficial and thoroughly enjoyed the immersive nature of VR. Furthermore, we observe how these interaction experiences influence participants' ability to transfer their pre-existing understanding of authentication into VR, resulting in a discrepancy in perceived security. Moreover, we identify users' conflicting expectations, encompassing their desire for an enjoyable VR experience alongside the assurance of secure VR authentication. Building upon our findings, we propose recommendations aimed at addressing these expectations and alleviating potential conflicts.
Distributing machine learning predictors enables the collection of large-scale datasets while leaving sensitive raw data at trustworthy sites. We introduce a learning technique that is scalable to a large number of users, satisfies Differential Privacy, and is applicable to non-trivial tasks, such as CIFAR-10. For a large number of participants, communication cost is one of the main challenges. We achieve a low communication cost by requiring only a single invocation of an efficient secure multiparty summation protocol. By relying on state-of-the-art feature extractors, we are able to utilize differentially private convex learners for non-trivial tasks such as CIFAR-10. Convex learners have proven to have a strong utility-private tradeoff. Our experimental results show that for $1{,}000$ users with $50$ data points each, our scheme outperforms state-of-the-art scalable distributed learning methods (differentially private federated learning, short DP-FL) while requiring around $500$ times fewer communication costs: For CIFAR-10, we achieve a classification accuracy of $67.3\,\%$ for an $\varepsilon = 0.59$ while DP-FL achieves $57.6\,\%$. We also show the learnability properties convergence and uniform stability.
Voice-controlled smart speaker devices have gained a foothold in many modern households. Their prevalence combined with their intrusion into core private spheres of life has motivated research on security and privacy intrusions, especially those performed by third-party applications used on such devices. In this work, we take a closer look at such third-party applications from a less pessimistic angle: we consider their potential to provide personalized and secure capabilities and investigate measures to authenticate users (``PIN'', ``Voice authentication'', ``Notification'', and presence of ``Nearby devices''). To this end, we asked 100 participants to evaluate 15 application categories and 51 apps with a wide range of functions. The central questions we explored focused on: users' preferences for security and personalization for different categories of apps; the preferred security and personalization measures for different apps; and the preferred frequency of the respective measure. After an initial pilot study, we focused primarily on 7 categories of apps for which security and personalization are reported to be important; those include the three crucial categories finance, bills, and shopping. We found that ``Voice authentication'', while not currently employed by the apps we studied, is a highly popular measure to achieve security and personalization. Many participants were open to exploring combinations of security measures to increase the protection of highly relevant apps. Here, the combination of ``PIN'' and ``Voice authentication'' was clearly the most desired one. This finding indicates systems that seamlessly combine ``Voice authentication'' with other measures might be a good candidate for future work.
Distributed differentially private learning techniques enable a large number of users to jointly learn a model without having to first centrally collect the training data. At the same time, neither the communication between the users nor the resulting model shall leak information about the training data. This kind of learning technique can be deployed to edge devices if it can be scaled up to a large number of users, particularly if the communication is reduced to a minimum: no interaction, i.e., each party only sends a single message. The best previously known methods are based on gradient averaging, which inherently requires many synchronization rounds. A promising non-interactive alternative to gradient averaging relies on so-called output perturbation: each user first locally finishes training and then submits its model for secure averaging without further synchronization. We analyze this paradigm, which we coin blind model averaging (BlindAvg), in the setting of convex and smooth empirical risk minimization (ERM) like a support vector machine (SVM). While the required noise scale is asymptotically the same as in the centralized setting, it is not well understood how close BlindAvg comes to centralized learning, i.e., its utility cost. We characterize and boost the privacy-utility tradeoff of BlindAvg with two contributions: First, we prove that BlindAvg converges towards the centralized setting for a sufficiently strong L2-regularization for a non-smooth SVM learner. Second, we introduce the novel differentially private convex and smooth ERM learner SoftmaxReg that has a better privacy-utility tradeoff than an SVM in a multi-class setting. We evaluate our findings on three datasets (CIFAR-10, CIFAR-100, and Federated EMNIST) and provide an ablation in an artificially extreme non-IID scenario.
Abstract For anonymous communication networks (ACNs), Das et al. recently confirmed a long-suspected trilemma result that ACNs cannot achieve strong anonymity, low latency overhead and low bandwidth overhead at the same time. Our paper emanates from the careful observation that their analysis does not include a relevant class of ACNs with what we call user coordination where users proactively work together towards improving their anonymity. We show that such protocols can achieve better anonymity than predicted by the above trilemma result. As the main contribution, we present a stronger impossibility result that includes all ACNs we are aware of. Along with our formal analysis, we provide intuitive interpretations and lessons learned. Finally, we demonstrate qualitatively stricter requirements for the Anytrust assumption (all but one protocol party is compromised) prevalent across ACNs.
Abstract Quantifying the privacy loss of a privacy-preserving mechanism on potentially sensitive data is a complex and well-researched topic; the de-facto standard for privacy measures are ε-differential privacy (DP) and its versatile relaxation (ε, δ)-approximate differential privacy (ADP). Recently, novel variants of (A)DP focused on giving tighter privacy bounds under continual observation. In this paper we unify many previous works via the privacy loss distribution (PLD) of a mechanism. We show that for non-adaptive mechanisms, the privacy loss under sequential composition undergoes a convolution and will converge to a Gauss distribution (the central limit theorem for DP). We derive several relevant insights: we can now characterize mechanisms by their privacy loss class, i.e., by the Gauss distribution to which their PLD converges, which allows us to give novel ADP bounds for mechanisms based on their privacy loss class; we derive exact analytical guarantees for the approximate randomized response mechanism and an exact analytical and closed formula for the Gauss mechanism, that, given ε, calculates δ, s.t., the mechanism is (ε, δ)-ADP (not an over-approximating bound).
Anonymous communication (AC) is a fundamental building block in numerous privacy enhancing technologies and applications. While there is a successful line of research and development on anonymous communication, it is an open question whether current approaches for anonymous communication networks are optimal. As a key step towards identifying how much anonymity an optimal AC network can provide, we take a different approach: we search for inherent limitations that apply to AC networks and outline the landscape of achievable anonymity. In a recent work [3], we have shown the first such upper bounds for anonymity, i.e., that certain combinations of bandwidth overhead, latency overhead and strong anonymity are impossible to achieve when faced with a global and passive network-level eavesdropper and node-level eavesdropper (i.e., a passive attacker). In this work, we show that the combination of secret-sharing and onion routing are able to escape our prior impossibility bounds, but we also prove novel impossibility bounds for such, more powerful protocols. For hybrid protocols that combine secret sharing and mix-nets techniques, the upper bounds on anonymity are significantly lower than for pure mix-nets, for the same latency and bandwidth overhead. In particular, while such hybrid protocols exhibit more resilience against compromisation than mix-net-like protocols, strong anonymity, low latency overhead and low bandwidth overhead still cannot be simultaneously achieved. Our work leaves as an open problem whether this combination of secret-sharing and onion routing is a theoretical effect or whether it can be realized in practice.
This work investigates the fundamental constraints of anonymous communication (AC) protocols. We analyze the relationship between bandwidth overhead, latency overhead, and sender anonymity or recipient anonymity against the global passive (network-level) adversary. We confirm the trilemma that an AC protocol can only achieve two out of the following three properties: strong anonymity (i.e., anonymity up to a negligible chance), low bandwidth overhead, and low latency overhead. We further study anonymity against a stronger global passive adversary that can additionally passively compromise some of the AC protocol nodes. For a given number of compromised nodes, we derive necessary constraints between bandwidth and latency overhead whose violation make it impossible for an AC protocol to achieve strong anonymity. We analyze prominent AC protocols from the literature and depict to which extent those satisfy our necessary constraints. Our fundamental necessary constraints offer a guideline not only for improving existing AC systems but also for designing novel AC protocols with non-traditional bandwidth and latency overhead choices.
Many applications, such as anonymous communication systems, privacy-enhancing database queries, or privacy-enhancing machine-learning methods, require robust guarantees under thousands and sometimes millions of observations. The notion of r-fold approximate differential privacy (ADP) offers a well-established framework with a precise characterization of the degree of privacy after r observations of an attacker. However, existing bounds for r-fold ADP are loose and, if used for estimating the required degree of noise for an application, can lead to over-cautious choices for perturbation randomness and thus to suboptimal utility or overly high costs. We present a numerical and widely applicable method for capturing the privacy loss of differentially private mechanisms under composition, which we call privacy buckets. With privacy buckets we compute provable upper and lower bounds for ADP for a given number of observations. We compare our bounds with state-of-the-art bounds for r-fold ADP, including Kairouz, Oh, and Viswanath's composition theorem (KOV), concentrated differential privacy and the moments accountant. While KOV proved optimal bounds for heterogeneous adaptive k-fold composition, we show that for concrete sequences of mechanisms tighter bounds can be derived by taking the mechanisms' structure into account. We compare previous bounds for the Laplace mechanism, the Gauss mechanism, for a timing leakage reduction mechanism, and for the stochastic gradient descent and we significantly improve over their results (except that we match the KOV bound for the Laplace mechanism, for which it seems tight). Our lower bounds almost meet our upper bounds, showing that no significantly tighter bounds are possible.
Many applications require robust guarantees against thousands and sometimes millions of observations, such as anonymous communication systems, privacy-enhancing database queries, or privacyenhancing machine-learning methods. The notion of r-fold Approximate Differential Privacy (ADP) offers a well-established framework with a precise characterization of the degree of privacy after r observations of an attacker. However, existing bounds for r-fold ADP are loose and, if used for estimating the required degree of noise for an application, can lead to overcautious choices for perturbation randomness and thus to suboptimal accuracy. We present a numerical (although widely applicable) method for capturing the privacy loss of differentially private mechanisms under composition, which we call privacy buckets. With privacy buckets we compute provable upper and lower bounds for ADP for a given number of observations. We compare our bounds with state-of-the-art bounds for r-fold ADP, including Kairouz, Oh, and Viswanath’s composition theorem (KOV), concentrated dfferential privacy and the moment’s accountant. We compare these bounds for the Laplace mechanism, the Gauss mechanism, for real-world timing leakage data and for the stochastic gradient descent and we significantly improve over their results (with the exception that the KOV bound seems tight for the Laplace mechanism). Moreover, our lower bounds almost meet our upper bounds, showing that no significantly tighter bounds are possible. *The authors are in alphabetical order. Both authors equally contributed to this work.
This technical report discusses three subtleties related to the widely used notion of differential privacy (DP). First, we discuss how the choice of a distinguisher influences the privacy notion and why we should always have a distinguisher if we consider approximate DP. Secondly, we draw a line between the very intuitive probabilistic differential privacy (with probability 1 − δ we have ε-DP) and the commonly used approximate differential privacy ((ε, δ)-DP). Finally we see that and why probabilistic differential privacy (and similar notions) are not complete under post-processing, which has significant implications for notions used in the literature. Differential privacy Differential privacy and its relaxation, approximate differential privacy (ADP) have been used widely in the literature. ADP classically quantifies, in terms of parameters ε and δ, the privacy of a mechanism that releases statistical information on databases with sensitive information in a privacy-preserving way. A prime example is releasing the answers to statistical queries on a database with medical records without revealing information about the individual records in the database. Moreover, differential privacy finds applications in machine learning, smart metering and even anonymous communication. Technically, differential privacy is often defined roughly like this: Definition 1 (Differential privacy). A mechanism M is (ε, δ)-differentially private, where ε ≥ 0 and δ ≥ 0, if for all neighboring databases D0 and D1, i.e., for databases differing in only one record, and for all sets S ⊆ [M ], where [M ] is the range of M , the following in-equation holds: Pr [M(D0) ∈ S] ≤ e · Pr [M(D1) ∈ S] + δ. We say that a mechanism satisfies pure differential privacy, if it satisfies (ε, 0)-differential privacy and we say that it satisfies approximate differential privacy (ADP), if it satisfies (ε, δ)-differential privacy for δ > 0. Differential privacy is quite helpful for defining privacy, particularly since it is preserved under arbitrary post-processing: no matter which functions are applied on the output of a differentially private mechanism, the results do not reduce the privacy guarantee in any way. More formally, if M is (ε, δ)-differentially private and T is any randomized algorithm, then T (M), defined as T (M)(x) = T (M(x)) is also (ε, δ)-differentially private. Background and further reading For background on the notions, definitions and applications of differential privacy we want to refer to the comprehensive book on differential privacy by Aaron Roth and Cynthia Dwork [2]. As a further read, particularly on the intuition behind the notions, we would recommend some of the great blog posts by Frank McSherry on what differential privacy is and isn’t about [6] and his illustrative take on the qualitative difference of ε and δ in (ε, δ)-differential privacy [7]. At this point we don’t want to delve into more intuition on why differential privacy is interesting or on the mathematical foundations behind the definition. Instead, we will simply claim that in a perfect world we would typically prefer to have pure (ε, 0)-differential privacy, as this would provide us with a clean and (more or less) easily understandable definition that provides us with deniability. In fact, if would make definitions of differential privacy a bit easier and a common pitfall could be avoided.
Many applications, such as anonymous communication systems, privacy-enhancing database queries, or privacy-enhancing machine-learning methods, require robust guarantees under thousands and sometimes millions of observations. The notion of r-fold approximate differential privacy (ADP) offers a well-established framework with a precise characterization of the degree of privacy after r observations of an attacker. However, existing bounds for r-fold ADP are loose and, if used for estimating the required degree of noise for an application, can lead to over-cautious choices for perturbation randomness and thus to suboptimal utility or overly high costs. We present a numerical and widely applicable method for capturing the privacy loss of differentially private mechanisms under composition, which we call privacy buckets. With privacy buckets we compute provable upper and lower bounds for ADP for a given number of observations. We compare our bounds with state-of-the-art bounds for r-fold ADP, including Kairouz, Oh, and Viswanath's composition theorem (KOV), concentrated differential privacy and the moments accountant. While KOV proved optimal bounds for heterogeneous adaptive k-fold composition, we show that for concrete sequences of mechanisms tighter bounds can be derived by taking the mechanisms' structure into account. We compare previous bounds for the Laplace mechanism, the Gauss mechanism, for a timing leakage reduction mechanism, and for the stochastic gradient descent and we significantly improve over their results (except that we match the KOV bound for the Laplace mechanism, for which it seems tight). Our lower bounds almost meet our upper bounds, showing that no significantly tighter bounds are possible.
Computing differential privacy guarantees is an important task for a wide variety of applications. The tighter the guarantees are, the more difficult it seems to be to compute them: naive bounds are simple additions, whereas modern composition theorems require either a search over a potentially confusing parameter space (for the composition theorem of Kairouz, Oh, and Visvanath), require computing many moments (for Rényi DP) or finding parameters of a function that limits the privacy loss (for concentrated DP). The best known approach for tight differential privacy bounds, called Privacy Buckets, provides requires running a fairly complex implementation of numerical approximations. In this work, we provide an easy-to-use interface for computing state-of-the-art differential privacy guarantees by simply accessing a website. Guarantees for the widely used Laplace mechanism and for the similarly popular Gauss mechanism can be computed by simply stating the scale parameter of the noise, the sensitivity, and the number of compositions. Privacy guarantees for more complex distributions can be computed by uploading two histograms. This work bridges the gap between the best known theoretical results for computing differential privacy guarantees and privacy analysts and users benefiting from such guarantees.
We present Loopix, a low-latency anonymous communication system that provides bi-directional 'third-party' sender and receiver anonymity and unobservability. Loopix leverages cover traffic and brief message delays to provide anonymity and achieve traffic analysis resistance, including against a global network adversary. Mixes and clients self-monitor the network via loops of traffic to provide protection against active attacks, and inject cover traffic to provide stronger anonymity and a measure of sender and receiver unobservability. Service providers mediate access in and out of a stratified network of Poisson mix nodes to facilitate accounting and off-line message reception, as well as to keep the number of links in the system low, and to concentrate cover traffic. We provide a theoretical analysis of the Poisson mixing strategy as well as an empirical evaluation of the anonymity provided by the protocol and a functional implementation that we analyze in terms of scalability by running it on AWS EC2. We show that a Loopix relay can handle upwards of 300 messages per second, at a small delay overhead of less than 1.5 ms on top of the delays introduced into messages to provide security. Overall message latency is in the order of seconds - which is low for a mix-system. Furthermore, many mix nodes can be securely added to a stratified topology to scale throughput without sacrificing anonymity.
Özgür Dagdelen合作论文数Department of Computer Science, Technische Universität Darmstadt1