Coronary Heart Disease (CHD) has remained one of the foremost causes of death in the world, and thus, there is a need to ensure that there are dependable early diagnosis mechanisms that would aid clinicians in making decisions at the right time. The rapid development of electronic health records and sensor-based medical data has presented more opportunities in predictive analytics in healthcare than ever before. However, the sensitivity, complexity, and scale of health data require robust analytical models and a safe and reliable data processing system. In this respect, machine learning (ML) methods have become effective instruments in deriving significant patterns of heterogeneous healthcare data. This study hypothesizes an ensemble learning framework that is used in the early identification of CHD. The proposed ensemble model is more accurate and stronger in predictions than any of the individual models by incorporating several ML classifiers. The study provides a scalable method to prevent cardiovascular diseases, and the model may help healthcare professionals to identify high-risk patients at an early stage and, thus, implement interventions in time and enhance patient outcomes. Experimental results demonstrate that the ensemble model outperforms conventional ML models, highlighting its effectiveness as a supportive diagnostic tool for CHD prediction.
Channel estimation in massive MIMO (Multiple-Input Multiple-Output) systems with one-bit ADCs (Analog-to-Digital Converters) faces severe quantization distortion that degrades conventional estimators and limits downstream detection and precoding. While cGAN (Conditional GAN)-based learning approaches can recover CSI (Channel State Information) from heavily quantized pilots, their susceptibility to adversarial perturbations has received limited attention despite the security sensitivity of wireless physical-layer processing. This paper presents the first framework that combines auxiliary classifier supervision with adversarial training for robust one-bit massive MIMO channel estimation. The proposed AC-GAN (Auxiliary Classifier GAN) discriminator is trained to jointly perform source discrimination (real/fake) and SNR (Signal-to-Noise Ratio) regime classification, encouraging disentangled, regime-aware feature learning that standard cGAN designs cannot achieve. We further integrate PGD (Projected Gradient Descent)-based adversarial training with curriculum scheduling into a minmax learning objective to improve robustness under white-box attacks. Simulations on a $\mathbf{6 4}$-antenna, 8-user uplink demonstrate that AC-GAN attains 15.3% lower NMSE (Normalized Mean Squared Error) than cGAN under clean conditions (-23.3 dB vs. $-20.8 \mathbf{~ d B}$) and achieves $\mathbf{7 8. 6 \%}$ robust accuracy under strong PGD attacks $(\varepsilon=0.1)$, outperforming adversarially trained cGAN $(37.4 \%)$ by 41.2 percentage points. The results show consistent gains across SNR regimes, pilot lengths, and attack types, including unseen attack strategies, supporting the use of auxiliary supervision as a practical and effective approach to more secure DL (Deep Learning)-enabled channel estimation.
Physical-layer security in sixth-generation (6G) networks faces a concrete threat that standard propagation-noise models overlook: adversarial perturbations injected during pilot transmission can catastrophically destabilise deep-learning-based channel estimators while remaining nearly imperceptible under standard power constraints. Existing countermeasures based on adversarial training or denoising autoencoders either demand prohibitive retraining overhead or generalise poorly to attack families unseen at training time. This paper addresses both shortcomings through Channel Estimation via Manifold Optimisation and Inference (CE-MOI), a denoising diffusion probabilistic model (DDPM)-driven framework that simultaneously restores accurate channel state information (CSI) and detects adversarial activity. The central insight is that legitimate channel matrices concentrate on a low-dimensional diffusion manifold; adversarial perturbations drive estimates off this manifold, producing a measurable reconstruction error $E_{\text{rec}}$ that serves as a principled detection statistic. CE-MOI refines corrupted initial estimates through iterative latent gradient descent constrained to the learned manifold and operates as a plug-and-play (PnP) wrapper around any existing least-squares (LS) or deep neural network (DNN) estimator, requiring no retraining. Evaluated in a multiuser multiple-input multiple-output (MU-MIMO) uplink with $M=32$ antennas and $U=4$ users across nine signal-to-noise ratio (SNR) levels and twelve perturbation conditions in simulation, CE-MOI confines normalised mean squared error (NMSE) degradation to within 1.65 dB of the clean baseline under strong white-box attacks, achieves area under the ROC curve (AUC) $>0.95$ for both white-box and black-box detection, and requires only $\approx 1.5$ GFLOPs and $\mathbf{0. 2 2 M}$ parameters, $1,928 \times$ fewer FLOPs than a score-based diffusion baseline at comparable accuracy.
As sixth-generation (6 G) networks transition toward edge-intelligence architectures, deep learning (DL)-based channel estimators face critical security vulnerabilities from adversarial perturbations. This paper investigates the adversarial robustness of a lightweight distilled channel estimator (D-CE) framework designed specifically for resource-constrained 6G edge devices. We systematically evaluate D-CE's resilience against four gradientbased adversarial attacks: Basic Iterative Method (BIM), Fast Gradient Sign Method (FGSM), Momentum Iterative Method (MIM), and Projected Gradient Descent (PGD). Experimental results demonstrate that defensive distillation achieves a 14-fold reduction in computational complexity while reducing attack success rate (ASR) by $\mathbf{2 5}-\mathbf{4 0 \%}$ compared to undefended baselines. The D-CE framework maintains normalized mean squared error (NMSE) within $\mathbf{1} \boldsymbol{-} \mathbf{2 ~ d B}$ of clean baseline performance under moderate adversarial perturbations, significantly outperforming convolutional neural networks (CNNs) that experience catastrophic estimation failures. These findings position defensive distillation as a practical and effective defense mechanism for secure channel state information (CSI) acquisition in adversarial 6G environments.
Sixth-generation (6 G) networks will increasingly rely on artificial intelligence (AI)/machine learning (ML) at the physical layer (PHY) to enable ultra-low latency, massive connectivity, and intelligent radio operation, but this shift also exposes new attack surfaces to adversarial manipulation. This survey synthesizes 30 recent studies (2021-2025) on adversarial threats and defenses in AI-enabled $\mathbf{6 G ~ P H Y}$, covering gradientbased attacks (e.g., Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD)) alongside spoofing, jamming, and man-in-the-middle (MitM) adversaries targeting millimeterwave (mmWave) beam prediction, channel estimation, reconfigurable intelligent surface (RIS), semantic communications, and Open Radio Access Network (O-RAN) intelligence. We organize defenses into adversarial training (up to 57.90 % robustness improvement), privacy-preserving federated learning (FL) (up to $\mathbf{1 7. 4 5 \%}$ defense gain), defensive distillation ($\mathbf{9 8. 3 \%}$ accuracy under attack), cryptographic/authentication mechanisms (up to 100 % verification accuracy), and reinforcement learning (RL)based adaptive strategies. While these approaches are effective in specific settings, open challenges remain in achieving crosstechnology generalization, mitigating heterogeneous concurrent attacks, and meeting strict real-time and resource constraints. The survey concludes with research directions toward unified, efficient, and deployable PHY defense frameworks for trustworthy 6 G.
Physical-layer security in sixth-generation (6G) multiple-input multiple-output (MIMO) systems demands channel estimators that are simultaneously accurate, adversarially resilient, and computationally tractablerequirements that existing methods address only in part. Gradient-based pilot attacks such as Projected Gradient Descent (PGD), Fast Gradient Sign Method (FGSM), and Carlini-Wagner (C&W) can degrade classical and deep-learning estimators by ${4}-{8} \text{d B}$ in normalised mean squared error (NMSE) while remaining near the thermal noise floor. This paper proposes a secure dual-denoising diffusion framework, termed Diffusion $-\mu$ CE, that addresses all three requirements through a single compact architecture. The key idea is to deploy the same distilled convolutional neural network (CNN) diffusion prior in two complementary roles: an input purifier that removes off-manifold adversarial components from the pilotinduced initial estimate, and an output refiner that suppresses residual artefacts after the main reverse trajectory. Every reverse step is augmented with a $\mu$-channel estimation data-consistency correction that re-anchors the trajectory to the observed pilots, preventing adversarial gradients from hijacking the sampling path. The student CNN is distilled from a high-capacity U-Net teacher with $2.1 \times 10^{6}$ parameters, reducing the deployed model to only $5.5 \times 10^{4}$ parameters and 5.5-22.4 Gfloating-point operations (FLOPs)80-100 times fewer than the score-based diffusion baseline. Evaluated under seven adversarial attacks and five natural noise types, Diffusion-$\mu$ CE confines NMSE degradation to at most 2.80 dB under any single attack, achieves a novel robustness gap of only $+0.38 \mathbf{~ d B}$ (versus $+3.16 \mathbf{~ d B}$ for the next-best diffusion method), and attains complete stochastic dominance over all seven baselines $(W=0, r_{\text{rb}}=1.000, p_{\text{Bonf}}=0.0034)$.
artificial intelligence (AI)-native sixth-generation (6G) multiple-input multiple-output (MIMO)-orthogonal frequency-division multiplexing (OFDM) receivers face two coupled impairments: gradient-aligned perturbations that exploit the sensitivity of deep learning (DL) estimators, and non-Gaussian noise that pushes pilot grids outside the distribution seen during training. This paper asks whether separating input restoration from estimator hardening can improve secure channel estimation (CE) under both impairment classes. We present Dual-Layer Defense System (DLDS), a modular two-layer pipeline for secure CE. Layer 1 uses a Wasserstein generative adversarial network with gradient penalty (WGANGP) generative prior and multi-restart latent-space descent to project corrupted inputs onto a learned clean-channel manifold. Layer 2 uses a compact student estimator distilled from a Projected Gradient Descent (PGD)-trained teacher convolutional neural network (CNN), producing the final robust channel state information (CSI) estimate at lower complexity than the teacher. In a simulation-based evaluation on DeepMIMO and CDL-C augmented MIMO-OFDM channels, DLDS achieves -18.5 dB clean normalised mean squared error (NMSE), 0.9 dB average degradation under diverse noise, and 4.8 dB degradation under strong attacks at $\epsilon=3.0$. DLDS attains $3.24 \times$ relative robustness over channel estimation deep neural network baseline (CE-DNN), with two-sided Welch tests giving $p<0.0001$ and Cohen's $d=3.82$. These results indicate that manifold purification and robust distillation provide complementary protection in the tested simulation setting, while adaptive end-to-end attacks and over-the-air validation remain outside the present scope.
The rise of cyberattacks has led to an increase in the creation of fake websites by attackers, who use these sites for advertising products, transmit malware, or steal valuable login credentials. Phishing, the act of soliciting sensitive information from users by masquerading as a trustworthy entity, is a common technique used by attackers to achieve their goals. Spoofed websites and email spoofing are often used in phishing attacks, with spoofed emails redirecting users to phishing websites in order to trick them into revealing their personal information. Traditional solutions for detecting phishing websites rely on signature-based approaches that are not effective in detecting newly created spoofed websites. To address this challenge, researchers have been exploring machine-learning methods for detecting phishing websites. In this paper, we suggest a new approach that combines the use of blacklists and machine learning techniques such that a variety of powerful features, including domain-based features, abnormal features, and abnormal features based on URLs, HTML, and JavaScript, to rank web pages and improve classification accuracy. Our experimental results show that using the proposed approach, the random forest classifier offers the best accuracy of 93%, with FPR and FNR as 0.12 and 0.02, with a Precision of 90%, Recall of 97% an F1 Score of 93%, and MCC of 0.85.
The increasing occurrence of assaults on Natural Language Processing (NLP) systems has sparked worry about their resilience and trustworthiness, particularly in crucial applications like sentiment analysis, machine translation and natural language deduction. To tackle these weaknesses this research presents GAN-Guarded Fine-Tuning (GAN-FT) an training structure that utilizes Generative Adversarial Networks (GANs) to bolster the resilience of extensive language models such, as BERT and RoBERTa. In contrast, training GAN with feature matching (GAN FW) uses a generative opponent to generate various and logically connected adversarial instances with less computational burden. Regular assessments on datasets such as IMDB, Yelp and SNLI indicate that GAN-FT consistently decreases the effectiveness of attacks while enhancing accuracy and adaptability across domains. This suggests its effectiveness in countering familiar and new word substitution forms of attack. This study highlights the impact of incorporating GANs into training methods and sheds light on understanding models better and making decisions effectively in this field of research. Though there are scalability and quick implementation challenges, in real-time scenarios, GAN-FT is a starting point for looking into adversarial defence mechanisms and broadening its use in more intricate NLP assignments. The research advances the reliability and security aspects of NLP systems, which is a move towards combatting adversarial risks in the ever-changing AI domain.
The safety and reliability of semantic segmentation networks face a major threat from adversarial perturbations which create problems in autonomous driving and medical imaging systems. The research develops a strong model-agnostic system which identifies and counteracts segmentation model attacks. The method applies uncertainty-based post-processing methods through pixel-wise entropy and dispersion metrics to detect adversarial inputs across different models without changing their internal structure. The method of adversarial robust knowledge distillation enables the transfer of defense capabilities from a high-capacity teacher network to an efficient student model which maintains segmentation accuracy and resource-limited deployment resilience. The integrated pipeline achieves state-of-the-art detection accuracy ($84.20 \%$) and segmentation robustness under various attack scenarios through empirical evaluation on standard benchmarks using convolutional and transformerbased segmentation architectures which outperform conventional baselines. The results demonstrate that robust distillation when used with uncertainty analysis leads to the development of reliable semantic segmentation networks for safety applications.
In recent years, diabetes has become more prevalent due to unhealthy lifestyles, obesity, and other factors. Diabetes is a chronic and dangerous disease due to its major complications that may lead to visual impairment, heart disease, death, kidney failure, and others. Therefore, the need for early detection of diabetes has increased in the health sector to prevent complications and other diseases that it leads to. With the advancement of technology, early prediction and detection of diseases have become a popular topic in scientific research. Machine learning has helped increase the effectiveness of early detection of diseases, as stated in our study on the early detection of diabetes. Six machine learning algorithms: XGBoost(XGB), Random Forest(RF), Decision Tree(DT), Support Vector Machine(SVM), Logistic Regression(LR), and K-Nearest Neighbors(KNN), were used and applied to the PIDD dataset. To improve the efficiency of the algorithms, we preprocessed the data, used the data standardization technique, and to reduce the noise and dimensions, principal component analysis (PCA)was used. Finally, after evaluating the performance of the algorithms according to several criteria, the best performance without PCA was in favor of SVM and LR with an accuracy of 78.65% and 78.39%, respectively. Still, LR was better than SVM when using PCA with an accuracy of 79.17%.
Malicious URL detection has become a critical aspect of cybersecurity, as malicious URLs, including phishing, malware distribution, and other malicious activities are increasingly prevalent. Traditional approaches, such as signature-based and rule-based methods, have shown limitations in adapting to evolving threats, highlighting the need for more advanced techniques. This study investigates the application of deep learning techniques, including CNN, LSTM, RNN, GRU, and Transformers, to malicious URL detection. The study explores the performance of these algorithms across different datasets, focusing on the accuracy, strengths, and weaknesses of each approach. This paper identifies key challenges, through a comprehensive analysis of recent studies. In addition, highlight emerging trends, innovations, and gaps in the literature, suggesting potential improvements for future research. The review also discusses future directions, including the development of hybrid models and stacking techniques to enhance accuracy while reducing overfitting. Our work aims to provide a clear understanding of modern deep learning approaches to malicious URL detection and provides important insights for researchers to build more robust and adaptive systems.
The phrase “data visualization” has taken on a lot of significance in today's world since it allows users of any level of expertise to comprehend data. Data visualization is a technique that turns a set of small and large raw data into visual data so that the user is capable of analyzing, comprehending, and discovering correlations, patterns, and trends from the data. This research study reviews some types of data visualization and their primary uses. In addition, there is a great need for data visualization in many fields, and this research focuses on the field of sales, where it uses two sales-related data sets to show the sales analysis and generates an interactive dashboard with a variety of data visualizations by using the Power BI tool, which is a service for business analytics and is extensively utilized for business intelligence, reporting, and data analysis across numerous industries.
This paper presents an intelligent phishing-detection system that combines lightweight transformer models with explainable AI for real-time protection across email and web browsing. Two fine-tuned DistilBERT classifiers are trained separately for (i) email text and (ii) URLs, using large public datasets (82.5k emails and 641k URLs) to ensure coverage and general usability. The system’s methodology integrates robust preprocessing, class-imbalance handling, and standard metrics (accuracy, precision, recall, F1, AUC-ROC), while LIME provides human-interpretable, token-level explanations. Experimental results show high performance in controlled tests for both email and URL models (0.99–1.00 across key metrics), with the email model marginally stronger overall; PR/ROC curves and confusion matrices confirm reliability and low error rates. The approach demonstrates that compact transformers can meet accuracy and latency targets for phishing detection while maintaining transparency via XAI, and it contributes a deployable prototype (Flask backend + browser extension) suitable for real-world integration and further work on adversarial robustness and continuous learning.
This paper explores and investigates the development and implementation of a deep learning-based privacy compliance framework for Internet of Things (IoT) applications. The geometric growth in potentially sensitive data from the vast generations of data as IoT devices proliferate raises critical security and privacy concerns for individuals and organisations. Additionally, rising cybersecurity attacks reshape how data is handled, especially with IoT devices storing sensitive client data. Simultaneously, GDPR and CCPA which are rigorous data protection regulations have created an urgent and imperative need for automated compliance solutions. Utilising these regulatory standards, this research harnesses deep learning methodologies, specifically Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks to tackle these complex challenges while employing public IoT datasets to test. Cutting-edge features are incorporated into the proposed deep-learning framework including threat characterising, anomaly detection, dynamic policy adjustment, and regulatory compliance checking/verification for IoT environments, with the additionality of a devised pipeline encompassing comprehensive data processing, model training, threat detection, compliance checking, and an intuitive user-friendly alert system. The findings concluded that the deep-learning models achieved high accuracy of 99% in threat detection, with general compliance achieving 91% and over of its IoT dataset (RT-IoT2022), concurrently while the integrated compliance checking functions effectively enforced privacy policies from a ‘docx’ file. This research contributes significantly to the relevant field by offering a novel, automated approach to handling privacy compliance in intricate IoT landscapes, prospectively reducing organisational resource expenditure while enhancing overall data protection measures.
Breast cancer continues to be one of the most prevalent and lethal cancers impacting women globally. Given the significant rates of occurrence and death linked to late diagnosis, it is essential to create precise and timely detection techniques. This research explores the use of machine learning algorithms to categorize tumor samples from the Breast Cancer Wisconsin (Original) dataset into benign or malignant types. A thorough preprocessing pipeline was developed, incorporating missing value treatment, feature scaling, and class balancing to improve data quality and model efficacy. Ten classifiers were assessed, which included Logistic Regression, Support Vector Machines (SVM), K-Nearest Neighbors (KNN), Random Forest, and Neural Networks. The findings highlighted marked enhancements in precision following optimization, emphasizing the efficacy of customized preprocessing methods. The research emphasizes the promise of machine learning for aiding in early breast cancer detection and suggests avenues for future studies centered on larger datasets, enhanced feature engineering, and interpretability.
With the great development of social media and people's increasing reliance on it, these platforms produce huge amounts of data daily, which reflect users' opinions and interactions in various fields. In this research, we chose Twitter for the study, as it is one of the most prominent platforms that allow opinions to be freely and directly expressed. Due to the huge volume of data published on Twitter, analyzing and understanding it becomes a difficult task without relying on advanced techniques such as deep learning (DL). Therefore, this study came to compare several popular methods in analyzing tweet sentiment while evaluating their performance using different criteria: accuracy, error rate, precision, recall, F1 score, and training time. In this study, we used the most common methods in sentiment analysis, namely Recurrent Neural Network (RNN), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and Bidirectional Encoder Representations from Transformers (BERT). The results showed that BERT is the best method for this task, achieving 91% accuracy and a 9% error rate, but its training time was much longer compared to other methods. Although BERT is the best choice, RNN or CNN are better alternatives when speed is a priority and resources are limited.
In the evolving field of cybersecurity, detecting malicious activity in high-dimensional network data remains a persistent challenge for traditional machine learning (ML) techniques. This study investigates the use of convolutional variational autoencoders (VAEs) to generate latent features that enhance the performance of various ML classifiers on the 2015 NSL-KDD dataset. Classifiers, including Gaussian Naïve Bayes (GNB), support vector machines (SVMs) with Radial Basis Function (RBF) kernel, decision trees, and dense neural networks, were evaluated using metrics such as accuracy, precision, recall, F1 score, and the Matthews Correlation Coefficient (MCC). To assess the effectiveness of VAEs, Principal Component Analysis (PCA) was used as a baseline dimensionality reduction method, and performance comparisons were made. The best-performing model was an SVM with an RBF kernel, a PCA (threshold = 0.92), and a VAE with six latent features, achieving an accuracy of 82.8%, an F1 score of 0.830, and an MCC of 0.682. The results indicate that VAEs can significantly enhance classifier performance, particularly in GNB and SVM models, suggesting their value in developing more effective intrusion detection systems. Received: 23 August 2024 | Revised: 3 June 2025 | Accepted: 15 July 2025 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement The data that support the findings of this study are openly available in NSL-KDD Dataset at https://github.com/t-taylor/cvaes-research/tree/main/results. Author Contribution Statement Thomas Taylor: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing – original draft, Writing – review & editing, Visualization. Amna Eleyan: Conceptualization, Methodology, Writing – original draft, Writing – review & editing, Supervision, Project administration. Mohammed Al-Khalidi: Validation, Resources, Writing – original draft, Writing – review & editing, Visualization, Supervision.
In this modern era, the identification of SMS spam is crucial due to the significant risk that spam poses to users. In this research, various supervised machine learning models (Support Vector Machine (SVM), Naive Bayes, and Random Forest) and transformer-based models (RoBERTa, DistilBERT) were utilized and trained on a relatively large, new dataset (super_sms_dataset) that was published in early 2024, which reached 67k records. In order to compare the performance of these models on different datasets, the same models were run on a different, smaller, and common dataset, the UCI dataset, which contains 5574 records. Consequently, transformer-based models outperformed traditional machine learning models, with the RoBERTa model achieving an impressive performance of 99.46% accuracy on the "super_sms_dataset" dataset.
Breast cancer is a prevalent disease affecting millions of women around the world. A key factor in improving the outcome of patients with breast cancer is early detection and classification. The use of convolutional neural networks (CNNs) has shown promising results for the analysis of various medical images, including the classification of breast cancer. This paper presents an overview of the breast cancer classification problem and demonstrates how a CNN can be effectively utilized for this task. Additionally, numerous papers have been presented and compared in terms of CNN structures, datasets, images, and accuracy. Different CNN models have been found to be effective at detecting breast cancer, which affects its accuracy. It should be recognized, however, that the accuracy of this algorithm depends on both the size of the dataset and the number of images that are used. As a result, it can be concluded that the number of images, datasets, or even the CNN approach can be used case-by-case to have higher accuracy. Finally, the results of accuracy should be expanded based on the analysis of one parameter in upcoming research. As soon as the best accuracy has been achieved, additional parameters may be added.