We extend a disclosure risk measure defined for population based frequency tables to sample based frequency tables. The disclosure risk measure is based on information theoretical expressions, such as entropy and conditional entropy, that reflect the properties of attribute disclosure. To estimate the disclosure risk of a sample based frequency table we need to take into account the underlying population and therefore need both the population and sample frequencies. However, population frequencies might not be known and therefore they must be estimated from the sample. We consider two probabilistic models, a log-linear model and a so-called Pólya urn model, to estimate the population frequencies. Numerical results suggest that the Pólya urn model may be a feasible alternative to the log-linear model for estimating population frequencies and the disclosure risk measure.
Statistical agencies are making increased use of the internet to disseminate census tabular outputs through web-based flexible table-generating servers that allow users to define and generate their own tables. The key questions in the development of these servers are: (1) what data should be used to generate the tables, and (2) what statistical disclosure control (SDC) method should be applied. To generate flexible tables, the server has to be able to measure the disclosure risk in the final output table, apply the SDC method and then iteratively reassess the disclosure risk. SDC methods may be applied either to the underlying data used to generate the tables and/or to the final output table that is generated from original data. Besides assessing disclosure risk, the server should provide a measure of data utility by comparing the perturbed table to the original table. In this article, we examine aspects of the design and development of a flexible table-generating server for census tables and demonstrate a disclosure risk-data utility analysis for comparing SDC methods. We propose measures for disclosure risk and data utility that are based on information theory.
Frequency tables disseminated by statistical agencies have always been of high interest. However, the agencies have to ensure that the risk of identifying individuals and disclosing individuals’ attributes from the released data is low. Therefore they assess the risk of disclosure and apply statistical disclosure control (SDC) methods if necessary. The main objective of this work is to measure dislosure risk in population based frequency tables. The disclosure risk assessment of such tables is often based on the so-called threshold rule. A cell of the table is of high disclosure risk according to this rule if the cell value does not exceed a certain threshold, for example 2. In this work we propose to measure the disclosure risk in an alternative way. Our approach takes the entire table (and rows/columns of the table) into consideration. We introduce a disclosure risk measure, which is based on information theoretical definitions, such as the entropy and the conditional entropy. There are two main types of SDC methods. Pre-tabular methods, such as record swapping, alter the values of a variable (or more variables) for selected individuals
Statistical agencies assess the risk of disclosure before releasing data. Unacceptably high disclosure risk will prevent a statistical agency from disseminating the data. The application of statistical disclosure control (SDC) methods aims to provide sufficient protection and make the data release possible. The disclosure risk of tabular data is typically quantified at the level of table cells. However, the evaluation of disclosure risk can require the assessment of the table as a whole, for example in the case of online flexible table generators. In this paper we use information theory to develop a disclosure risk measure for population-based frequency tables. The proposed disclosure risk measure quantifies the risk of attribute disclosure before and after an SDC method is applied. The new measure is compared to alternative disclosure risk measures developed at the Office for National Statistics.