Current big data providers offer little-to-no control over how your data is used once it is collected. Data cooperatives are an alternative to these companies and give control of personal data back to the data providers (whether they be people or organizations), allowing them to determine which of their data is used and how their data is used. Data cooperatives can serve as a more ethical alternative to other big data solutions, and have already seen success in the real world. However, supporting software must be developed to ensure the privacy of data providers beyond cooperative promises. In this paper, we expand upon our previous work applying homomorphic encryption (HE) to secure the personally identifiable information (PII) of data providers in data cooperatives that use graph storage. Data cooperatives are expected to store and query over data of varying security levels, including PII, low-security (where anonymization alone is sufficient), and public domain information. To facilitate graph storage, we introduce a multidimensional graph storage technique designed specifically for data cooperatives that mix cleartext, encrypted, and anonymized heterogeneous edges over a heterogeneous set of vertices. We demonstrate a HE query watchdog, which prevents incidental data leakage at query runtime and prior to decryption when proper rules are provided. This watchdog is complementary to existing work preventing data leakage prior to query runtime. This watchdog's operations are dominated by any reasonably-complex query.
A handful of companies currently hold large collections of data about most people. In addition to the questionable ethics of collecting personal data with few-to-no options to limit what these companies collect, there exist exceptionally few ways to regulate how your data is stored and used once it is collected. Furthermore, these data collections cannot be easily cross-referenced to gain insight. Data cooperatives provide an alternative to these separated collections of data. As a participant-driven organization, similar to a credit union, data cooperatives have a vested interest in preserving the privacy of individuals while offering insight similar to other big data analytics. Another bonus of the data cooperative model is the voluntary (and ethical) sourcing of data. The downside of giving participants the freedom to choose which data they contribute is incomplete data sets. To help address this, we adapt label propagation, a semi-supervised learning algorithm for community detection based on partially labeled data, to work over homomorphically encrypted (HE) graphs. We also adapt triangle counting and a vertex scoring scheme to work over directed heterogeneous-vertex, heterogeneous-edge HE graph data.
k-Anonymity is widely used to preserve privacy of individuals by ensuring at least k data points are indistinguishable from one-another. In this paper, we apply k-anonymity to our federated heterogeneous graph storage technique, category cluster. This application of k-anonymity allows select graph edges that would otherwise have to be encrypted with homomorphic encryption to remain in cleartext. We also demonstrate k-core anonymization, a graph anonymization technique based on k-core decomposition. We benchmark anonymized versions of Youtube and LiveJournal social networks against their original counterparts. This anonymization method preserves core membership well (less than 0.7% induced error). k-core anonymization also preserves the distribution of longest paths; however, our results show this distribution is shifted in phase from the original.
The future quantum network repeater is envisioned to primarily serve the role of creating entanglement between nodes and distilling those entanglements to an optimal level of performance. During our investigation, we implemented a multi-pass protocol for entanglement distillation and tested it on the IBM-Q environment, demonstrating successively improved results after multiple passes. We implemented two versions of multi-pass distillation, BBPSSW and DEJMPS, with a focus on optimizing the use of qubits, via the reset-and-reuse capability of the IBM implementation. The novel feature of reset-and-reuse can be a game-changer and can minimize the number of qubits required for large-scale applications. We also found that, though it is currently not possible to implement a criterion for continued distillation passes as a run-time feedback loop, the process can be studied through post-circuit data analysis. Our results also show that fidelity alone may guide us to discard some approaches that show success based on other metrics, such as entanglement success and success of transmitting a bit of data. The fidelity was experimentally found to be excessively low, for this complex process of multi-pass distillation.
"Big data" continues to grow in influence with few competitors able to challenge them. In order to slow the growth of and eventually replace these "data silos", we must enable competition from alternative sources that respect users' privacy, such as data cooperatives. In our previous work, we proposed an architecture for a privacy-preserving data cooperative that relies on homomorphic encryption (HE) to ensure data privacy and demonstrated ring-based BFS, degree centrality, and farness centrality over HE graph data. In this paper we expand our suite of HE graph algorithms to include single-source shortest-path, all-pairs shortest-path, minimum spanning tree, harmonic centrality, random walk, and betweenness centrality over HE graph data. These graph analysis algorithms support the core service of a data cooperative: to provide data and insights (or aggregates) to the service of the cooperative's clients (researchers, companies, governments, etc.) while maintaining the privacy of their users.
Data such as an individual's income, favorite sports team, typical commute route, vehicle maintenance history, medical records, etc. are typically not useful for making large-scale decisions such as where to build a new hospital, identifying which roads are in need of upkeep, and the like.However, aggregates of of these data across hundreds of individuals are useful to governments and to companies.Data cooperatives/unions offer a place for individuals to store their data and a service of data aggregation and interpretation to governments, non-profit organizations, and businesses while maintaining individuals' anonymity.We propose the use of anonymization techniques coupled with graph algorithms over homomorphically encrypted (HE) graphs as a basis of analysis for this accumulated data.We believe this approach ensures individuals' privacy and anonymity while preserving the usefulness of the plaintext data.
Ram Dantu合作论文数Department of Computer Science & Engineering University of North Texas7