
Maintaining access control policies is an ongoing process to ensure required but not excessive authorizations. Organizations thus leverage various data sources to ease this maintenance. Among these data sources are access control matrices, attributes, access logs, and transaction logs. While research reasonably covers the former data sources, the potential of transaction logs remains untapped. We pave the way for transaction logs as a data source in access control by (i) expressing them with a formalization, (ii) pinpointing them in typical Identity and Access Management (IAM) infrastructures, and (iii) grounding them in IAM processes. We conclude that access control transaction logs are valuable data sources for improving analytical capabilities for IAM.
The rapid growth of decentralized systems in the Web3 ecosystem has introduced numerous challenges, particularly in ensuring data security, privacy, and scalability [3, 8]. These systems rely heavily on distributed architectures, requiring robust mechanisms to manage data and interactions among participants securely. One critical aspect of decentralized systems is key management, which is essential for encrypting files, securing database segments, and enabling private transactions. However, securely managing cryptographic keys in a distributed environment poses significant risks, especially when nodes in the network can be compromised [9]. This research proposes a decentralized database scheme specifically designed for secure and private key management. Our approach ensures that cryptographic keys are not stored explicitly at any location, preventing their discovery even if an attacker gains control of multiple nodes. Instead of traditional storage, keys are encoded and distributed using the BFLUT (Bloom Filter for Private Look-Up Tables) algorithm [7], which enables secure retrieval without direct exposure. The system leverages OrbitDB [4], IPFS [1], and IPNS [10] for decentralized data management, providing robust support for consistency, scalability, and simultaneous updates. By combining these technologies, our scheme enhances both security and privacy while maintaining high performance and reliability. Our findings demonstrate the system's capability to securely manage keys, prevent unauthorized access, and ensure privacy, making it a foundational solution for Web3 applications requiring decentralized security.
The notion that federated learning ensures privacy simply by keeping data local is widely acknowledged to be flawed. Cryptographic techniques such as Multi-Party Computation (MPC) and Fully Homomorphic Encryption (FHE) address this issue by concealing the model during the training procedure, but their extreme computational and communication overhead makes them impractical for real-world deployment. However, we argue that such strong guarantees are unnecessary. Even with full-model encryption, black-box attacks remain possible during the prediction phase, since model outputs are eventually revealed to the querier. This suggests that instead of enforcing perfect privacy during training, it is sufficient to ensure that the leakage during training is no higher than the leakage during prediction. To achieve this, we generalize POSEIDON (NDSS 2021), a state-of-the-art FHE-based federated learning approach, by selectively encrypting only the components of the model necessary to match the privacy level of the prediction phase. Our method identifies the parts of the model that contribute most to information leakage and prioritizes their encryption, significantly reducing computational and communication overhead. Our experiments on dense neural networks show that encrypting only the last layer is often sufficient to hinder white-box attacks, improving efficiency by a linear factor in the number of layers. For deeper models, multiple layers may require encryption, but our approach still achieves a substantial speedup compared to full-model encryption.
Any organization providing services to users typically requires the users to share some degree of personal information in order to access these services. However, users seldom have complete visibility or control over their own data, justifiably raising serious privacy concerns. Additionally, organizations are often inhibited from collaboration due to the increased risk of privacy breaches, and corresponding reporting requirements. Increasing privacy focused legislation worldwide such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) create complex protection and reporting requirements that need to be systematically enforced. The Right To Be Forgotten (RTBF) recommendations included in GDPR essentially demand that every individual has the right to request for erasure of data pertaining to it, and thus in effect be forgotten from the system on demand. We propose a novel approach called CRISP (Consensus-enabled, Redactable, Immutable, Securely shareable, Provable) that uses permissioned enterprise blockchains for trustless interoperation among a network of identifiable organizations providing the ability to support RTBF. It combines off-chain data storage and distributed resource sharing with blockchain consensus mechanisms and decentralized access control. We describe the various components of CRISP and delineate the detailed steps of its operation. Results of an extensive set of experiments clearly establish the viability of our approach.
In modern browsers, interactive 3D graphics are enabled by the WebGL component, which also serves as a vector for browser fingerprinting. Browser fingerprinting offers the benefit of allowing web application service providers to capture data about a web user's browsing activity and as such establish user authenticity. One drawback of fingerprinting, however, is that the data collected can reveal unique (private) details about a user's browsing behaviours which is in contravention of digital privacy laws such as the GDPR. To address this issue, anti-browser fingerprinting solutions such as randomising or blocking WebGL parameters have been proposed. However, these solutions face challenges with detectability and web compatibility, often resulting in performance degradation and poor user experience. In this paper, we propose a performance-efficient anti-browser fingerprinting mechanism for WebGL that is robust to detectability and offers improved user experience. We achieve this by extending the JShelter browser extension by: (1) generating realistic and valid WebGL parameters, reducing the likelihood of detection while preserving the functionality of visited websites; and (2) enhancing the spoofing (randomisation) mechanism in JShelter, to ensure that the spoofed values are indistinguishable from those of real hardware configurations. The results of our empirical study indicate reduced WebGL related errors in JShelter from over 150 to zero. Additionally, the modified JShelter version transfers data more efficiently and achieves faster speeds compared to the unmodified version, thus improving privacy, usability, and compatibility.
The data economy thrives on data-centric collaboration between organizations. However, open data sharing remains a pipe dream without addressing pragmatic, regulatory, and strategic concerns – which include data protection and confidentiality. Generative artificial intelligence supports the production of realistic synthetic datasets on demand, and is a promising technology to alleviate such concerns. However, the replacement of an original dataset with a synthetic dataset incurs a specific trade-off between utility and privacy, and the appropriateness of this trade-off is highly context- and application-dependent. Manually establishing and managing different synthetic generators that have diverging properties is error-prone and time-consuming, lacks flexibility, and thus is costly and impractical. This paper introduces Data Chameleon, a novel self-adaptive data management architecture for different synthetic data generators. Data Chameleon adaptively samples from different synthetic data generators in function of the data request at hand. Furthermore, in a longer-term adaptation loop, Data Chameleon monitors and evaluates the overall suitability of the available generators, to monitor evolutions in data demand, or possible concept drifts. Based on this, the Chameleon autonomously decides to re-train existing generators, or instantiate additional ones in a self-adaptive manner. The Data Chameleon architecture enhances the practical applicability of synthetic data generation, enabling more efficient and secure data sharing in real-world scenarios.
This study examines metadata-assisted detection of supply chain attacks in Infrastructure as Code (IaC), focusing on metadata's role in identifying security smells. Metadata, including dependency relationships and author records, provides insights into IaC scripts but remains underutilized by detection tools. The evaluation of static IaC smell detection tools highlights their limitations in incorporating metadata analysis. To address this, a methodology integrating metadata and dependency analysis was developed to identify security smells in dependency chains. An analysis of 482 Ansible Galaxy repositories identified vulnerabilities in 45 dependency chains, including reliance on deprecated dependencies (CWE-477), hard-coded credentials (CWE-798), and improper file permissions (CWE-280). Additionally, three repositories contained security vulnerabilities associated with output (CVE-2024-8775) and logging (CVE-2017-7550). The findings highlight the necessity of integrating metadata analysis with static code analysis for detecting security smells. This approach enhances IaC security and mitigates risks related to supply chain attacks.
We have conducted a comprehensive examination of the fee payment process within the Dash CoinJoin implementation. We have discovered that the fee payment process leaks privacy information. This privacy leakage leads to the potential linking of addresses that the CoinJoin transaction design attempts to maintain private. We demonstrate that the transactions involved in the mixing process are identifiable and provide classification rules to build a transaction data set with appropriate classifications. Using this dataset, we highlight how the fee payment mechanism reveals sensitive information about users' mixing activities. We achieve this by linking the fee-payment transactions and analyzing their timestamps. Finally, our analysis reveals a significant occurrence of address reuse within the fee payment process and we provide a summary of the statistics obtained.
Modern networks are composed of a diverse group of devices and applications, all of which speak different protocols and exhibit varied network behaviors. Understanding these communication patterns is a critical requirement for a network security analyst to enforce effective access control against malicious behavior. The heterogeneity of devices, the diverse communication patterns and the lack of detailed documentation makes it harder to detect malicious behavior. Towards this end, we model the network communication patterns akin to natural language and design a custom transformer architecture to provide the necessary comprehension. We use natural language processing (NLP) transformers for the analysis of network traffic flows, especially for those flows exhibiting high spatial contextualization and temporal correlation. Through extensive experiments, we demonstrate that our model provides a reasonably generic understanding of network flows when applied to solve critical network problems such as application identification, device-type fingerprinting and threat identification. We tested our approach on three diverse data sets with the following results: (a) IoT device-type fingerprinting, an average recall of 97
The revised eIDAS regulation (eIDAS 2.0) advocates a shift back to user control over digital credentials, introducing the European Digital Identity Wallet. This shift aims to enhance privacy by allowing citizens to disclose personal data in a controlled manner selectively. As the keys to which the credentials are bound must be stored securely, a secure storage mechanism is essential-one that is not only secure but also accessible through the available technology stack and compliant with eIDAS 2.0. In support of the European Digital Identity Wallet, the EU Commission published an Architecture and Reference Framework together with a set of Implementing Acts to ensure interoperable solutions. However, the current versions only identify a high-level set of requirements and do not provide insights on satisfying them through actionable implementations. Secure storage is a crucial aspect that remains inadequately addressed, highlighting the need for comprehensive security and privacy guidelines to ensure a robust solution. To address this gap, we provide a threat model explicitly designed for the secure storage component of the wallet. This allows for identifying potential threats and a set of effective controls to secure the implementations and serves as a practical tool to assist architects in making informed decisions when selecting an implementation that best meets their system's security and privacy requirements. In addition, it reinforces essential assurance activities, such as certification, testing, and attestation required by the eIDAS 2.0 to maintain a trusted state for secure storage.
Organizations operate in a constantly changing environment. They face many disruptions that can affect their sustainability. Cyber resilience is a de facto concept on which they can rely on to ensure their sustainability. It is defined as the ability of a nation, organization, or mission or business process to anticipate, withstand, recover from, and evolve to improve capabilities in the face of adverse conditions, stresses, or attacks on the supporting cyber resources it needs to function. It is influenced by several disciplines. From these disciplines, it is necessary to extract the requirements that the organization must comply with to be able to be cyber resilient. Once the requirements have been extracted, they need to be cleaned up to eliminate duplicates, so that only a single set of requirements applicable to the organization remains. Manual extraction is tedious. We propose an approach integrating Machine Learning to automate the extraction of cyber resilience requirements from documents written in natural language. This extraction is followed by the identification of duplicates to retain single set of requirements that will be implemented for cyber resilience enhancement.
Timed Data Release (TDR) is a practical security mechanism designed to safeguard data until a specified time has elapsed. Recently, a more general paradigm, referred to as Event-driven Data Release (EDR) has gained attention. EDR enables data to be released based on the occurrence of a specified event. Existing blockchain-based constructions for EDR either neglect privacy considerations or restrict themselves to specialized message spaces (e.g., signatures) while achieving provable privacy guarantees. Despite its practical relevance, constructing a privacy-preserving EDR framework that supports general message spaces on blockchain with formal security guarantees remains an open challenge. In this paper, we address this gap by introducing P-EDR, a privacy-preserving event-driven data release framework for general message spaces, built on smart contracts.P-EDR enables a data sender to encrypt its data, which can be securely released upon the occurrence of specified events. P-EDR incorporates a novel, tailored signature-based witness encryption, SGWE, designed to encrypt data messages within a general message space, contingent upon the presence of a valid signature of a referencing message. Using SGWE, decryption of ciphertext is only possible with a valid signature. We formally define the desired security model and present an efficient construction of SGWE. Additionally, we propose a practical design for P-EDR and we rigorously prove the security guarantees of our construction. Our experimental results for evaluating the performance of the proposed construction demonstrate the effectiveness of our approach.
Ensuring Real-Time Bidding compliance with GDPR and the ePrivacy Directive remains challenging. Existing solutions, like the IAB's Transparency and Consent Framework and OpenRTB, lack transparency and fail to secure legally valid user consent. We propose a blockchain-based framework that decentralizes consent management, giving users direct control. The system automates consent handling, provides immutable compliance proof, and includes a smartphone app for users to manage consent per website. Additionally, we introduce a specificity value for IAB Audience Taxonomy elements, helping users assess the privacy impact of sharing data. By enhancing autonomy, transparency, and accountability, our approach strengthens trust in programmatic advertising while maintaining GDPR compliance without disrupting the industry.
In this paper, we attack the problem of storing purpose and consent metadata. This special kind of data brings a constraint where a false positive signifies the violation of a regulation. Our approach allows necessary metadata for purpose and consent maintenance and enforcement to be efficiently organized as a custom filter data structure to enhance metadata storage and fast retrieval. Contrary to other filters in the literature, we show that our approach guarantees that no false positives occur to the users' consent. We analyze the configuration knobs of this data structure, discuss improvements, and our experimental results show that our method outperforms a more traditional approach of storing a foreign key to refer to a purpose storage table.
Investigating the attack process followed by automated scripts, tools or botnets and their threat actors is a widely researched area in the cyber security. There has also been emphasis on differentiating attacks performed by humans considering their adaptability which poses a greater threat. Limited number of features such as slower typing speed and typing mistakes in commands issued by human attackers on the compromised systems have been used to detect their presence. This paper presents a study of human attackers by deploying 15 honeypots in five locations worldwide, collecting and analysing attack data for two months. We propose a comprehensive feature set based on characteristics and patterns of issued commands and the usage of alphanumeric, modifiers, cursor and other keys in the attack process. We used these features to distinguish human attackers interacting with honeypots. Moreover, five case studies are discussed to provide insights into actions performed in the attack process. The results show various actions performed by human attackers ranging from executing basic commands for getting device information to more advanced actions such as downloading files, running scripts and removing traces of their activities.
Reusing open-source software (OSS) code has become standard in software development. When vulnerabilities are discovered in reused code, maintainers typically apply security patches. However, these patches often include non-vulnerability-related changes, such as code refactoring or updating a setting file. Applying a patch without distinguishing these changes can lead to unintended software malfunctions. Existing techniques do not account for non-remediation code lines in security patches. This study aims to mitigate unexpected failures caused by indiscriminate patch application. We propose a method for identifying the code lines that directly remediate vulnerabilities in security patches. By leveraging lexical preprocessing and Large Language Models (LLMs), our approach semantically classifies code lines within a security patch, distinguishing vulnerability-fixing changes from unrelated changes. In an experimental evaluation using security patches for 25 distinct vulnerability types, the proposed method achieved an F1 score of 0.88, improving 0.22 over the baseline. The results also indicate that classification accuracy decreases for vulnerability types requiring extensive modifications, such as injection and authentication vulnerabilities. Additionally, we revealed that nine out of twenty-five security patches (36
Privacy policies play a critical role in disclosing the collection and sharing of personal information. However, due to their complex and lengthy nature, users often find it difficult to comprehend and therefore tend to simply ignore them. This paper presents a novel approach that leverages the capabilities of large language models to aid users with automatic tools that enable them to discern the policies and assess the privacy risks while minimizing the reliance on human-labeled data. In particular, our approach automatically maps each paragraph of a privacy policy to predefined categories and extract the privacy attributes and their relationships with the first party collection and third party sharing. These attributes and relationships are represented as a graph, enabling the use of readily available graph databases such as Neo4j. This lends itself to answer privacy questions as Cypher queries and to identify inconsistent privacy statements.
A variant of Matrix Operation for Randomization or Encryption (v-MORE) is a lightweight, real number-based, fully homomorphic encryption scheme used for outsourced secure medical data processing. This paper shows the vulnerability of v-MORE to a ciphertext-only attack when it is used in privacy-preserving applications where the data is not random. From the ciphertexts of v-MORE, the attacker can solve for its secrets based on linear algebra and predict the plaintext from a new ciphertext with 100
This paper explores the application of the DUCA (Data Usage Control and Compliance Architecture) framework for privacy management in Cyber-Physical Systems (CPS). DUCA integrates Privacy-by-Design (PbD) principles, Privacy-Enhancing Technologies (PETs), and context-aware policy enforcement to support regulatory compliance and protect data throughout its lifecycle. A key focus of this work is the integration of Secure Elements (SEs)-including Trusted Execution Environments (TEE), Trusted Platform Modules (TPM), and Intel SGX-to enable privacy protection during data processing, complementing traditional safeguards for data at rest and in transit. The framework also supports emerging standards such as DICE and MARS to facilitate scalable trust management in heterogeneous CPS environments. We present DUCA's modular architecture and evaluate its applicability across representative use cases, including smart grids, eHealth, and AI-enabled infrastructures, demonstrating its effectiveness in enforcing privacy without compromising functionality.
We consider the problem of enforcing corporate governance control relying on cloud-based services. Extending previous work, we focus in particular on the support of delegation of the director privileges, enabling their dynamic and temporary assignment to a vice-director. Like previous work, our control relies on encrypted tags, which are here extended addressing the challenges introduced by dynamic delegation which operates on a time dimension orthogonal to the corporate governance control process. Our solution enables delegation while ensuring a vice-director to enjoy the director privileges only when delegation is active and not to operate as director for operations the vice-director has processed as employee (separation of duties). Our tag construction ensures integrity of the dynamic delegation control and protection against tag tampering.