The synthesis of Large Language Models (LLMs) with the Internet of Cloud (IoC) ecosystems creates multiple opportunities across diverse domains such as healthcare, finance, and smart cities. This study explores the combination of these technologies, focusing on their ability to profile authors and the associated privacy challenges. Two interesting experiments were conducted using real data from the Blog Authorship Corpus and the Reddit Self-Reported Depression Diagnosis (RSDD). Then the capabilities of two well-known LLMs (ChatGPT-4o and Llama 3-70B) were evaluated regarding the identification of sensitive demographic data of users such as gender, age, profession, and psychological conditions. Our findings highlight privacy risks in IoC environments, where user-authored logs, commands, and reports are already stored and analyzed by cloud-based LLM services. Moreover, key insights indicate that while LLMs improve precision and adaptability in textual data analysis, they also increase the potential risks of profiling detection in sensitive contexts. This research highlights the urgency of implementing robust privacy-preserving strategies to mitigate ethical risks and social impacts. Finally, it presents all the useful findings from the two experiments and then provides a detailed analysis of the results by comparing the two LLMs used.
The proliferation of Large Language Model (LLM)-based digital assistants has introduced significant privacy risks, as user queries containing Personally Identifiable Information (PII) are routinely transmitted to external AI services without adequate filtering or anonymization. This paper presents the implementation and evaluation of Controlled Query Routing (CQR), a privacy-preserving methodology designed to govern how user queries are processed in hybrid AI systems. Building upon the architectural framework proposed in our previous work, CQR integrates four sequential stages: PII detection using Named Entity Recognition (NER), risk scoring, routing decision, and anonymization with GDPR-aligned logging. The methodology was implemented as a working prototype—a security and privacy awareness digital assistant—deployed using FastAPI, Microsoft Presidio, and the Claude API. Evaluation across nine query scenarios demonstrates that CQR correctly identifies sensitive entities, assigns appropriate risk scores, and routes queries without exposing personal data to external models. The results confirm the feasibility of combining privacy-by-design principles with LLM-based assistants while supporting compliance with the EU General Data Protection Regulation (GDPR) and the AI Act.
In recent years, machine learning algorithms are increasingly dependent on large volumes of data for their training, including personal data, while at the same time the law has strengthened the right of individuals to have such data deleted, thus creating an inherent tension. Regulations such as the General Data Protection Regulation (GDPR) oblige an organization to erase personal data on request, but deleting a record from a database is not enough. A trained model retains the influence of that record in its parameters and may still expose it, for example, through membership inference. Machine unlearning has emerged in order to remove this influence from the model itself, and it has rapidly developed into an active research area. However, existing surveys have not provided a unified, verifiability-centered account of what is required to demonstrate that unlearning has actually occurred. This review provides a unified treatment of machine unlearning, beginning with the taxonomy of exact and approximate algorithms and the trade-off between efficacy, fidelity, and efficiency that governs them. It then examines the role of unlearning in privacy protection and its dual role in security, where it serves as a defense against poisoning and backdoors but also becomes an attack surface. Particular attention is given to evaluation, because the empirical tests of the literature can measure a removal but cannot prove it. On this basis, the review examines verifiable, federated, and decentralized unlearning, including the Proof of Unlearning and zero-knowledge constructions. Taken together, the review’s findings indicate that most methods assert rather than prove removal, while verifiable unlearning in federated and decentralized environments remains a central open problem.
Gamified applications are widely used across domains such as education, healthcare, and the workplace, offering a way to improve user engagement and motivation. However, their collection and processing of personal data pose significant privacy challenges. This paper examines how gamified systems manage user data and proposes a structured approach to support privacy-risk analysis by mapping system modules to the data elements they process and examining their structural interrelations. To this end, we examine the elements and data that, if compromised or misused, can harm users’ privacy. Building on this analysis, we explain how data elements connect to privacy risks and derive practical safeguards for privacy-aware design. The findings suggest that while gamification can increase user motivation and engagement, it simultaneously creates additional opportunities for privacy violations, especially where authentication, profiling, analytics, and interaction features are tightly coupled. The study concludes by highlighting the importance of integrating privacy-by-design principles into the development of gamified systems and discussing the limitations of a connectivity-based assessment.
Federated learning is gaining increasing traction, including in healthcare applications. The platform presented in this paper, developed by a multidisciplinary consortium, enables privacy-preserving training of machine learning models to generate predictions for patients with chronic obstructive pulmonary disease and comorbidities. In addition, data synchronization and monitoring are facilitated via the HL7 FHIR standard. The platform includes two front ends: a patient-facing smartphone app and a dashboard designed for healthcare professionals, currently in use at three hospitals in Italy, Estonia, and the Netherlands. Source code, synthetic datasets and fitted ML models will be released and indexed on zenodo.org . Initial ML results obtained from models trained with the platform are discussed. The overall architecture and its implementation in European hospitals is shown in this paper.
This study addresses the growing complexity of privacy protection in cloud computing environments (CCEs) by introducing a comprehensive socio-technical framework for self-adaptive privacy, complemented by an AI-driven beta tool designed for social media platforms. The framework’s three-stage structure—social, technical, and infrastructural—integrates context-aware privacy controls, dynamic risk assessments, and scalable implementation strategies. Key benefits include enhanced user-centric privacy management through customizable group settings and adaptive controls that respect diverse social identities. The beta tool operationalizes these features via a profile store for structured preference management and a recommendation engine that delivers real-time, AI-powered privacy suggestions tailored to individual contexts. Additionally, the tool’s safety scoring system (0–100) empowers developers and guides them in designing effective privacy solutions and mitigating risks. By bridging social context awareness with technical and infrastructural innovation, this framework significantly improves privacy adaptability, regulatory compliance, and user empowerment in CCEs. It provides a robust foundation for developing scalable and responsive privacy solutions tailored to evolving user needs.
With the increasing complexity of network infrastructures, anomaly detection and Quality of Service (QoS) assurance have become critical in maintaining secure and efficient operations. This paper identifies various machine learning algorithms that contribute to solving these challenges by analyzing network traffic, detecting anomalies, and optimizing QoS. Supervised learning algorithms, such as Support Vector Machines (SVMs), Random Forest, and K-Nearest Neighbors (KNN), have been effectively applied to real-time traffic classification and intrusion detection. These methods excel in identifying patterns that distinguish between legitimate and malicious network activity. Unsupervised techniques, such as K- means clustering, play a crucial role in detecting evolving anomalies without requiring labeled datasets, making them highly adaptable to dynamic environments. Additionally, ensemble learning methods, such as AdaBoost and XGBoost, have shown enhanced performance in both anomaly detection and QoS management through gradient boosting and regularization techniques. This paper explores these algorithms’ capabilities, providing insight into their application in real-world network scenarios and the benefits they bring to anomaly detection and QoS optimization.
The rapid expansion of Large Language Models (LLMs) within Internet of Cloud (IoC) ecosystems creates significant risks regarding data privacy, security, and compliance. Additionally, although LLMs support real-time decision making and intelligent cloud services, their use within IoC ecosystems may expose sensitive data to privacy risks due to their complex design. This paper explores how ten of the most recognized privacy challenges such as: unauthorized data access, model inversion, and data leakage, arise during the deployment of 20 commonly used LLMs in IoC ecosystems. It begins by outlining each privacy challenge, then explains its specific impact on IoC ecosystems, regulatory compliance, and severity levels. Moreover, this study introduces a comparative matrix that evaluates each LLM's level of compliance with these challenges. The matrix identifies which models meet privacy expectations and which do not, and includes examples of non-compliance, offering a clearer understanding of how these models differ in their exposure, vulnerability, and mitigation practices. The analysis reveals severe discrepancies across models, with many lacking sufficient transparency, effective consent management, and secure data deletion mechanisms. Finally, the findings emphasize the urgent need for a comprehensive privacy-by-design strategy and AI alignment protocols tailored to cloud-based LLM deployments.
To protect patient privacy while enabling proper research and data sharing in the healthcare environment, anonymization of health data is crucial. There are significant privacy, ethics, and legality concerns as a result of the continually expanding potential for re-identification and data manipulation brought about by the digitization and interconnectedness of healthcare data. InviseeAI addresses all these challenges by integrating sophisticated methods like secure multiparty computation, customized education, and differential privacy with traditional privacy techniques in a novel, AI-driven method towards data anonymization. InviseeAI in contrast to traditional alternatives provides a very extensible, open, platform for secure sharing of health data along with effective discovery of proper balance between the utility of data and privacy. In accordance with early findings, InviseeAI successfully reduces reidentification risk without compromising the analytical value of data sets so that enhanced partnerships in epidemiology, clinical studies, and operational health intelligence can be enabled.
Anonymization Techniques are among the main approaches for keeping people’s information private in the digital world, where the security of confidential data is paramount. While k-anonymity, l-diversity, and t-closeness have traditionally been applied to avoid direct re-identification through anonymization, these approaches cannot help avoid indirect re-identification using auxiliary data. Complementary to these are mathematical frameworks injecting controlled noise, such as differential privacy (DP), to protect data against reasoning attacks. This work has aimed to integrate DP within state-of-the-art anonymization techniques and has given case examples showing improved privacy without loss of data utility. Such a detailed inquiry into the advantages and disadvantages of these hybrid techniques will emphasize the possible utilization in business sectors like healthcare, banking, and location-based services.
The proliferation of Internet of Things (IoT) devices has brought tremendous convenience in our daily lives but has also brought significant privacy concerns. In recent years, many solutions have been found in the literature to address these challenges through advanced technologies such as Artificial Intelligence (AI). This paper aims to provide a comprehensive survey of the current landscape of IoT privacy, focusing on the role of AI in enhancing privacy measures. We categorize critical privacy challenges, outline AI strategies to address these challenges, and present AI-driven solutions that have shown real and substantial results in major sectors. We examine various AI techniques, assess their effectiveness, and highlight existing research gaps to inform future researchers. Our main contributions include a taxonomy of AI applications for IoT privacy, an analysis of AI-driven privacy solutions, and a discussion on the ethical implications and compliance requirements. This paper is recommended to researchers, practitioners, and policymakers seeking to develop secure and privacy-aware IoT systems. Unlike previous surveys that analyze thoroughly individual privacy-preserving methods, this study provides a multi layer synthesis of AI techniques tailored to IoT architectures and deployment realities, presenting a taxonomy grounded in both theoretical robustness and implementation feasibility.
Privacy design in cloud systems remains complex, with unclear processes and a mismatch between privacy engineering and cloud integration challenges. Developers play a pivotal yet underexplored role in this landscape. This study investigates developers' perspectives on privacy, focusing on self-adaptive privacy in cloud environments. Through six(6) qualitative interviews with developers from Greece, Spain, and the UK, the study uncovers valuable insights into their challenges and perspectives, contributing to the establishment of actionable privacy goals and a taxonomy of self-adaptive privacy requirements. The findings underscore the need for clearer guidance and actionable insights for developers to enhance privacy practices in cloud development.
The vast adoption of cloud computing has led to a new content in relation to privacy and security. Personal information is no longer as safe as we think and can be altered. In addition, Cloud Service Providers (CSPs) are still looking for new ways to raise the level of trust in order to gain popularity and increase their number of users. In this paper, a systematic literature review was carried out to identify the different methodologies, models and frameworks regarding privacy engineering and trust in cloud computing. A detailed review is produced on the specific area to bring forward all the work that has been carried out the recent years using a methodology with a number of different steps and criteria. Based on the findings from the literature review, we present the state-of-the-art on privacy and trust methodologies in cloud computing and we discuss the existing conventional tools that can assist software designers and developers.
Create, Read, Update and Delete operations (CRUD) are a well-established abstraction to model data access in software systems of different architectures. Most system requirements, generated during the specification phase, will be realized by combining these operations on different entities of the system under development. The majority of these requirements will be business operations and objectives. Security requirements come on top of business requirements in a mostly network-connected world and risk the existence of a software system as a business. Through the enforcement of privacy laws, modern systems must also legally comply with privacy requirements or face the possibility of high fines. While there is a great interest in methodologies to elicit security and privacy requirements, little has been done to practically apply those requirements during the software development phase. This paper investigates the implication of those four basic operations regarding security and privacy principles as they are implied by the law. Analysis findings aim to raise awareness among developers about privacy when implementing high-level business requirements, and result in a bottom-up compliance procedure regarding privacy and the GDPR by proposing a systematic approach in this direction.