This paper offers a comprehensive analysis of the statistical disclosure limitation (SDL) methodologies employed by the U.S. Census Bureau on the 2010 and 2020 Decennial Census releases under the perspective of the disclosure risk of the most vulnerable respondents. We first review the SDL methodology used up to the Decennial Census 2010, which was based on targeted swapping. Second, we examine recently reported reconstruction and reidentification results on the Decennial Census 2010 outputs, which form the foundation for the U.S. Census Bureau’s decision to switch to a differentially private (DP) method for the 2020 release. Third, we examine the actual privacy and data accuracy achieved by the DP method and compare with the privacy and accuracy offered by the formerly employed swapping mechanism. We conclude that the DP method is not an adequate solution to protect the typically sparse tables present in the Decennial Censuses because it does not offer meaningful privacy guarantees in general, it poorly protects the privacy of the most vulnerable respondents in particular, and it significantly degrades the quality of the released data. We also argue that the claimed disclosure risks of previous Census releases were overstated because of a flawed reidentification procedure. Therefore, the U.S. Census Bureau’s decision to change the SDL methodology to a DP-based one for the 2020 release was not only unwarranted, but it also reduced privacy and data quality compared to former releases.
Most methods in the literature on synthetic microdata (individual records) generation are parametric, that is, they require knowing or estimating the joint or the conditional distribution of the original microdata. This may be a significant hurdle unless the original microdata are multivariate normal. We propose a rank-based approach to generating synthetic microdata based on the permutation paradigm. We present three different methods and we analyze the utility and the confidentiality they afford. The third method is actually an extension of the second method that adds $k$ -anonymity protection against reidentification to the confidentiality against attribute disclosure offered by the first two methods. Our algorithms only require the identification of the marginal distributions of attributes and yield synthetic attributes that replicate the relationships between the original attributes exclusively based on ranks. This proposal is especially attractive for non-normal or multi-type microdata.
Several authors have claimed the “failure of anonymization,” despite over 50 years of research. We review privacy leaks reported over the past decades and conclude they were due to nonexistent or inadequate anonymization, rather than a lack of robust anonymization methods.
Synthetic data generation is a promising approach for sharing data for secondary purposes in sensitive sectors. However, to meet ethical standards and legislative requirements, it is necessary to demonstrate that the privacy of the individuals upon which the synthetic records are based is adequately protected. Through an expert consensus process, we developed a framework for privacy evaluation in synthetic data. The most commonly used metrics measure similarity between real and synthetic data and are assumed to capture identity disclosure. Our findings indicate that they lack precise interpretation and should be avoided. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information. The framework provides recommendations to effectively measure these types of disclosures, which also apply to differentially private synthetic data if the privacy budget is not close to zero. We further present future research opportunities to support widespread adoption of synthetic data.
Due to its small size and lifelong optical transparency, the fish Danionella cerebrum is an emerging model organism in biomedical research. How can this small vertebrate under 12 mm length produce sounds over 140 dB? We found that it possesses ...Motion is the basis of nearly all animal behavior. Evolution has led to some extraordinary specializations of propulsion mechanisms among invertebrates, including the mandibles of the dracula ant and the claw of the pistol shrimp. In contrast, vertebrate ...
The threat of reconstruction attacks has led the U.S. Census Bureau (USCB) to replace in the Decennial Census 2020 the traditional statistical disclosure limitation based on rank swapping with one based on differential privacy (DP), leading to substantial accuracy loss of released statistics. Yet, it has been argued that, if many different reconstructions are compatible with the released statistics, most of them do not correspond to actual original data, which protects against respondent reidentification. Recently, a new attack has been proposed, which incorporates the confidence that a reconstructed record was in the original data. The alleged risk of disclosure entailed by such confidence-ranked reconstruction has renewed the interest of the USCB to use DP-based solutions. To forestall a potential accuracy loss in future releases, we show that the proposed reconstruction is neither effective as a reconstruction method nor conducive to disclosure as claimed by its authors. Specifically, we report empirical results showing the proposed ranking cannot guide reidentification or attribute disclosure attacks, and hence fails to warrant the utility sacrifice entailed by the use of DP to release census statistical data.
In 2017, the United States Census Bureau announced that because of high disclosure risk in the methodology (data swapping) used to produce tabular data for the 2010 census, a different protection mechanism based on differential privacy would be used for the 2020 census. While there have been many studies evaluating the result of this change, there has been no rigorous examination of disclosure risk claims resulting from the released 2010 tabular data. In this study we perform such an evaluation. We show that the procedures used to evaluate disclosure risk are unreliable and resulted in inflated disclosure risk. Demonstration data products released using the new procedure were also shown to have poor utility. However, since the Census Bureau had already committed to a different procedure, they had no option except to escalate their commitment. The result of such escalation is that the 2020 tabular data release offers neither privacy nor accuracy.
In our article “Database Reconstruction Is Not So Easy and Is Different from Reidentification”, we show that reconstruction can be averted by properly using traditional statistical disclosure control (SDC) techniques, also sometimes called legacy statistical disclosure limitation (SDL) techniques. Furthermore, we also point out that, even if reconstruction can be performed, it does not imply reidentification. Hence, the risk of reconstruction does not seem to warrant replacing traditional SDC techniques with differential privacy (DP) based protection. In “Legacy Statistical Disclosure Limitation Techniques Were Not an Option for the 2020 US Census of Population and Housing”, by Simson Garfinkel, the author insists that the 2020 Census move to DP was justified. In our view, this latter article contains some misconceptions that we identify and discuss in some detail below. Consequently, we stand by the arguments given in “Database Reconstruction Is Not So Easy:: :”.
We review the use of differential privacy (DP) for privacy protection in machine learning (ML). We show that, driven by the aim of preserving the accuracy of the learned models, DP-based ML implementations are so loose that they do not offer the ex ante privacy guarantees of DP. Instead, what they deliver is basically noise addition similar to the traditional (and often criticized) statistical disclosure control approach. Due to the lack of formal privacy guarantees, the actual level of privacy offered must be experimentally assessed ex post , which is done very seldom. In this respect, we present empirical results showing that standard anti-overfitting techniques in ML can achieve a better utility/privacy/efficiency tradeoff than DP.
In our article "Database Reconstruction Is Not So Easy and Is Different from Reidentification", we show that reconstruction can be averted by properly using traditional statistical disclosure control (SDC) techniques, also sometimes called legacy statistical disclosure limitation (SDL) techniques. Furthermore, we also point out that, even if reconstruction can be performed, it does not imply reidentification. Hence, the risk of reconstruction does not seem to warrant replacing traditional SDC techniques with differential privacy (DP) based protection. In "Legacy Statistical Disclosure Limitation Techniques Were Not an Option for the 2020 US Census of Population and Housing", by Simson Garfinkel, the author insists that the 2020 Census move to DP was justified. In our view, this latter article contains some misconceptions that we identify and discuss in some detail below. Consequently, we stand by the arguments given in "Database Reconstruction Is Not So Easy:: :".
In recent years, it has been claimed that releasing accurate statistical information on a database is likely to allow its complete reconstruction. Differential privacy has been suggested as the appropriate methodology to prevent these attacks. These claims have recently been taken very seriously by the U.S. Census Bureau and led them to adopt differential privacy for releasing U.S. Census data. This in turn has caused consternation among users of the Census data due to the lack of accuracy of the protected outputs. It has also brought legal action against the U.S. Department of Commerce. In this paper, we trace the origins of the claim that releasing information on a database automatically makes it vulnerable to being exposed by reconstruction attacks and we show that this claim is, in fact, incorrect. We also show that reconstruction can be averted by properly using traditional statistical disclosure control (SDC) techniques. We further show that the geographic level at which exact counts are released is even more relevant to protection than the actual SDC method employed. Finally, we caution against confusing reconstruction and reidentification: using the quality of reconstruction as a metric of reidentification results in exaggerated reidentification risk figures.
Providing access to synthetic micro-data in place of confidential data to protect the privacy of participants is common practice. For the synthetic data to be useful for analysis, it is necessary that the density function of the synthetic data closely approximate the confidential data. Hence, accurately estimating the density function based on sample micro-data is important. Existing kernel-based, copula-based, and machine learning methods of joint density estimation may not be viable. Applying the multivariate moments' problem to sample-based density estimation has long been considered impractical due to the computational complexity and intractability of optimal parameter selection of the density estimate when the true joint density function is unknown. This paper introduces a generalised form of the sample moment-based density estimate, which can be used to estimate joint density functions when only the information of empirical moments is available. We demonstrate optimal parametrisation of the moment-based density estimate based solely on sample data by employing a computational strategy for parameter selection. We compare the performance of the moment-based estimate to that of existing non-parametric and parametric density estimation methods. The results show that using empirical moments can provide a reasonable, robust non-parametric approximation of a joint density function that is comparable to existing non-parametric methods. We provide an example of synthetic data generation from the moment-based density estimate and show that the resulting synthetic data provides a reasonable disclosure-protected alternative for public release.
Recent analysis by researchers at the U.S. Census Bureau claims that by reconstructing the tabular data released from the 2010 Census, it is possible to reconstruct the original data and, using an accurate external data file with identity, reidentify 179 million respondents (approximately 58% of the population). This study shows that there are a practically infinite number of possible reconstructions, and each reconstruction leads to assigning a different identity to the respondents in the reconstructed data. The results reported by the Census Bureau researchers are based on just one of these infinite possible reconstructions and is easily refuted by an alternate reconstruction. Without definitive proof that the reconstruction is unique, or at the very least, that most reconstructions lead to the assignment of the same identity to the same respondent, claims of confirmed reidentification are highly suspect and easily refuted.
Anonymization for privacy-preserving data publishing, also known as statistical disclosure control (SDC), can be viewed under the lens of the permutation model. According to this model, any SDC method for individual data records is functionally equivalent to a permutation step plus a noise addition step, where the noise added is marginal, in the sense that it does not alter ranks. Here, we propose metrics to quantify the data confidentiality and utility achieved by SDC methods based on the permutation model. We distinguish two privacy notions: in our work, anonymity refers to subjects and hence mainly to protection against record re-identification, whereas confidentiality refers to the protection afforded to attribute values against attribute disclosure. Thus, our confidentiality metrics are useful even if using a privacy model ensuring an anonymity level ex ante. The utility metric is a general-purpose metric that can be conveniently traded off against the confidentiality metrics, because all of them are bounded between 0 and 1. As an application, we compare the utility-confidentiality trade-offs achieved by several anonymization approaches, including privacy models (k-anonymity and $\epsilon$-differential privacy) as well as SDC methods (additive noise, multiplicative noise and synthetic data) used without privacy models.
Differential privacy (DP) is a privacy model that was designed for interactive queries to databases. Its use has then been extended to other data release formats, including microdata. In this paper we show that setting a certain ϵ in DP does not determine the confidentiality offered by DP microdata, let alone their utility. Confidentiality refers to the difficulty of correctly matching original and anonymized data, and utility refers to anonymized data preserving the correlation structure of original data. Specifically, we present two methods for generating ϵ -differentially private microdata. One of them creates DP synthetic microdata from noise-added covariances. The other relies on adding noise to the cumulative distribution function. We present empirical work that compares the two new methods with DP microdata generation via prior microaggregation. The comparison is in terms of several confidentiality and utility metrics. Our experimental results indicate that different methods to enforce ϵ -DP lead to very different utility and confidentiality levels. Both confidentiality and utility seem rather dependent on the amount of permutation performed by the particular SDC method used to enforce DP. Thus suggests that DP is not a good privacy model for microdata releases.
Methods for Privacy-Preserving Data Publishing (PPDP) have been recently shown to be equivalent to essentially performing some permutations of the original data. This insight, called the permutation paradigm, establishes a common ground uponwhich anymethod can be evaluated ex-post, but can also be viewed as a general ex-ante method in itself, where data are anonymized with the injection of suitable permutation matrices. It remains to develop around this paradigm a formal privacy model based on permutation. Such model should be sufficiently intuitive to allow non-experts to understand what it really entails for privacy to permute, in the same way that the privacy principles lying behind k-anonymity and differential privacy can be grasp by most. Moreover, similarly to differential privacy thismodel should ideally exhibit simple composition properties, which are highly handy in practice. Based on these requirements, this paper proposes a new privacy model for PPDP called Pα,β -privacy. Using for benchmark a one-time pad, an absolutely secure encryption method, this model conveys a reasonably intuitive meaning of the privacy guarantees brought by permutation, can be used ex-ante or ex-post, and exhibits simple composition properties. We illustrate the application of this new model using an empirical example.
Generating synthetic data for the dissemination of individual information in a privacy-preserving way is an approach that is often presented as superior to other statistical disclosure control techniques. The reason for such claim is straightforward at first glance: since all records disseminated are synthetic and not actual observed values, no individual can reasonably claim to face a privacy threat. Thus, and if the synthesizer used is good enough, synthetic data will potentially always offer a high level of information with low disclosure risk attached. Building on recent advances in the literature regarding the conceptualization of an intruder, this paper aims at challenging this claim by reassessing the privacy guarantees of synthetic data. Using the concept of a maximum-knowledge intruder, we demonstrate that synthetic data can in fact be always expressed as a re-arrangement of the original data and that, as a result, they may lead to configurations where disclosure risk may be higher than for non-synthetic disclosure control approaches. We illustrate the application of these results by an empirical example.
Yan-Xia Lin合作论文数School of Mathematics and Applied Statistics
University of Wollongong2
K. El Emam合作论文数University of Ottawa1