Data privacy has become a critical concern in recent years, particularly with the rising incidence of data breaches in sensitive domains such as healthcare. According to recent reports, the number of healthcare data breaches has increased by over 50
Over the past few years, the proliferation of deepfake content—digitally manipulated media that successfully alters appearances or voices—has created considerable challenges in social, ethical, and cybersecurity realms. Existing deepfake detection approaches, based mostly on Convolutional Neural Networks (CNNs), tend to have difficulty with capturing fine details of subtle facial irregularities and spatial relationships. To overcome this, we introduce a hybrid model that incorporates Graph Convolutional Networks (GCNs) and CNNs. Our model utilizes GCNs to study facial landmarks as graphs, preserving relational information, while CNNs concentrate on pixel-level features in images. By combining results from both models, we provide a strong methodology that takes advantage of the respective strengths. We test our approach on a deepfake dataset taken from kaggle, demonstrating better accuracy compared to conventional CNN-based methods, especially in detecting subtle manipulations. This work provides a novel hybrid system to improve the accuracy of deepfake detection.
Continuous collection of data from Global Navigation Satellite Systems (GNSS) is essential for assessing the internal dynamic processes of the Earth and for maintaining a reliable reference system for geodesy. Due to the sheer volume and degree of noise present in GNSS network data, it becomes impractical to identify unusual behavior using manual or semiautomated techniques. This research provides a hybrid deep learning framework to perform unsupervised anomaly detection in geodetic time-series data. The method utilizes a station-specific Long Short-Term Memory (LSTM) autoencoder to determine the typical time-series behavior of each individual GNSS station, and a physically motivated velocity constraint to ensure that anomalies found through the framework correspond to instantaneous and temporary kinematic deviations from the normal behavior of the station rather than being due to ambient noise. A unique component of this study is the development of a comprehensively validated methodology that utilizes synthetically generated datasets that mimic many real-world types of geodetic anomalies. To optimize the detection thresholds to achieve maximal F1-scores, the method was experimentally evaluated using three commonly utilized baseline methods: a Moving Average filter, Isolation Forest, and one-class SVM with Radial Basis Function kernel. Results obtained by testing the proposed framework using synthetic datasets showed that the hybrid framework consistently outperformed all baselines, achieving good recall and an average F1-score of 0.9914 across the evaluated stations, with particular success in precision-recall characterization compared to other approaches. Additionally, the proposed framework successfully identified the majority of previously documented events of anomalous behavior during historical validation, and maintained a low false positive rate. The results indicate that the proposed hybrid framework will be effective for fully automated largescale monitoring of GNSS networks, providing direct relevance to assessing resilient infrastructure and ensuring long-term geodetic stability.
In medical image diagnostics, classifying brain tumors appropriately would result in better outputs on the treatment end. So, this paper suggests an innovative approach for brain tumor image classification through federated learning and transfer learning to ensure privacy of data within the medical institutes. This methodology applies a pre-trained model, VGG-16 with ImageNet, and transfers the learned model, using less of the highly labeled data. The federated framework that uses the FedAvg algorithm for model aggregation ensures to keep data decentralized and private, thus enabling collaborative building of models without sharing actual data. It conducted its experiment on a dataset spread over 10 clients of brain tumors and encompassing a variety of medical centers. It had four classes-glioma, meningioma, notumor, and pituitary. The model performs good classification performance with the aid of a macro-averaged F1-score of 0.97. It shows the possibility of real-world health care, such a high-accuracy, decentralized AI models paving the way for privacy-focused collaborative advances in medical diagnostics.
The exponential growth of surveillance video data has surpassed the capabilities of manual monitoring, necessitating the development of intelligent systems for real-time suspicious behavior recognition (RTSBR). A spatio-temporal framework is proposed in this study, which utilizes Convolutional Neural Networks (CNNs) for spatial feature extraction and Long Short-Term Memory (LSTM) units for effective temporal sequence modeling, integrated within a Long-term Recurrent Convolutional Network (LRCN) architecture. Enhanced by multi-head self-attention mechanisms, the system captures context-aware features critical for detecting subtle behavioral anomalies. To address computational efficiency and scalability, the model incorporates optimized hyperparameters, bidirectional LSTMs, and feature fusion strategies, enabling robust performance across varying environments. Evaluation on benchmark datasets demonstrates superior accuracy—96.8 % on Avenue and 94.6% on VIRAT—outperforming existing approaches. Performance metrics, along with performance evaluation, validate the model’s suitability for real-time deployment in surveillance scenarios. Inference-time analysis confirms its readiness for latency-sensitive applications. Model pruning and quantization further reduce computational overhead, facilitating deployment on edge-based surveillance systems. The proposed RTSBR system delivers high scalability, adaptability, and detection accuracy, positioning it as a viable solution for proactive threat identification in complex multi-camera environments. This approach contributes to advancing real-time video analytics for intelligent security monitoring under dynamic operational constraints.
A fire can spread rapidly and unpredictably, causing significant damage, financial loss, and threats to human safety. Therefore, an effective fre detection solution is essential for preventing such hazards. Thus, a hybrid ConvNeXtV2 Ghost-based model (YOLO-GV2), built on the latest YOLOv10 framework, has been proposed to enhance the accuracy and efficiency in real-time scenarios. The ConvNeXtV2 serves as the backbone, strengthening the feature extraction and representation, thereby improving fre discrimination in visually complex environments. Additionally, the Ghost Convolution and C3Ghost modules are incorporated in the neck to reduce computational redundancy while maintaining a strong detection capability thereby making the model lightweight and suitable for resource-constrained devices. The experimental evaluation on a fre indoor dataset demonstrates that YOLO-GV2 outperforms state-of-the-art models, including Faster R-CNN and YOLOv10s. Compared to YOLOv10s, it achieves a 6.9% increase in precision, 10.1% improvement in recall, and gains of 7.7% and 9.4% in mean Average Precision (mAP) at Intersection over Union (IoU) threshold 0.50 (mAP@50) and across IoU thresholds from 0.50 to 0.95 (mAP@50–95), respectively. These results confirm that YOLO-GV2 provides a balanced trade-off between accuracy and efficiency, enabling robust real-time fire detection for safety-critical applications.
Quantification of Terrestrial Water storage (TWS) is hampered by an inherent observational gap: satellite gravimetry (GRACE/GRACE-FO) is able to give out mass data at the coarse monthly and spatial scales of pixel sizes of less than 300 km, whereas in situ piezometers are not able to cover the entire continent. Global Navigation Satellite System (GNSS) stations are also stations that measure the action of mass loading on the crustal elastic deformation at a high frequency, and it is a possible source of high-frequency monitoring. Non-hydrological multipath and noise effects are however notorious in extracting hydrological signals out of the GNSS vertical displacement. The idea we would like to present is a Physics-Informed Transformer (PIT) framework which tries to repurpose single-station GNSS series as a high-frequency hydrological sensor. In contrast to a typical deep learning architecture model, which directly maps meteorological inputs to displacement, our structure includes a so-called Kinematic Constraint Layer, which imposes the elastic loading equation, i.e. the equation of displacement, as the result of K elastic modulus multiplied by the resultant changes in weight, K. In the case of NAUS station in the Amazon Basin, the model automatically converges at an elastic admittance factor of K = −0.2425 mm/mm which compares with the theoretical Green functions. The framework attains Pearson correlation of R = 0.74 with the detrended GNSS readings but does a very good job in smoothing out high-frequency noise that afflicts the unconstrained black-box models. This method empirically reduces the GRACE-scale of information to daily resolution and thereby monitors short term bursts of floods that would not have been detected otherwise.
Automatic image captioning aims to acquire the learning of the correlation between the visual and the linguistic cognitive domain through the generation of descriptive captions with respect to input images. But it is affected by the semantic distance between image features and text, and problems in modeling fine-grained object interactions and contexts. We present a new encoder-decoder architecture, that utilizes optimal vision transformer + attention to the transformer by using cosine distance in order to get more context by providing better mapping. It relies on Inception-V3 to acquire fine-grained visual features and skip-gram features of the textual semantics. Another algorithm named improved walrus optimization dynamically updates the key hyperparameters to enhance convergence and performance. Experiments on the MS-COCO and Flickr8k datasets show that it is better than prior methods with BLEU-4 scores of 0.8864 and 0.8920, and higher performance gains in ROUGE, METEOR and CIDEr scores. In conclusion, the method outputs with better contextual details and flow of natural language, semantic fidelity. It is optimal in practice due to its cosine attention and adaptive tuning.
The increasing number of automobiles on the road has worsened traffic, raised questions about safety, and had an adverse effect on the environment. Traffic Signal Control (TSC) has become a key strategy to address these issues. While conventional TSC systems help manage urban traffic, their reliance on fixed schedules and simplified models limits their effectiveness in complex, real-world conditions. In order to overcome the above shortcomings, this paper introduces an Improved Twin-Actor Twin-Delayed Deep Deterministic Policy Gradient (ITATD3) algorithm designed for traffic signal control at single intersections. The main innovation of the ITATD3 proposed is in two aspects: (1) the incorporation of a spatial occupancy-based reward which in combination with waiting time, better measures road-space utilization and congestion than traditional metrics, and (2) the application of the Cheetah Optimization Algorithm (COA) for fine-tuning of key hyper parameters of the ITATD3 model to improve convergence speed and stability. The performance of the proposed model is compared against four benchmark models, like TATD3, TD3, D3QN, and DQN. Results show that ITATD3 performs consistently better than other models on various performance metrics, such as travel delay, average waiting time, traffic conflicts, spatial occupancy, throughput, and average travel time. The proposed ITATD3 reduced travel delay by a maximum of 68% and improved traffic throughput significantly compared to conventional models. These results reveal the benefits of ITATD3 as a viable solution for real-time data-driven traffic signal control in intelligent transportation systems.
The latest advancements in speech synthesis and voice conversion systems have created highly realistic synthetic voices, known as audio deepfakes. These developments pose significant risks to automatic speaker verification, authentication, and forensic investigations. Current detection methods often struggle with poor interpretability, overfitting, and limited generalization in real-world situations. To tackle these challenges, this paper introduces a new hybrid deepfake detection framework. This framework combines interpretable handcrafted acoustic features with PCA-compressed WavLM embeddings. This approach allows for a better understanding of low-level spectral artifacts and high-level contextual cues. Unlike previous hybrid or self-supervised detectors that depend on complex ensembles or unclear transformer back-ends, our design enhances robustness, transparency, and computational efficiency within a lightweight 106-D fused representation. Handcrafted features capture low-level acoustic properties like Mel-Frequency Cepstral Coefficients (MFCCs), spectral centroid bandwidth, Zero-Crossing Rate (ZCR), chroma, and Root-Mean-Square Energy (RMSE). Meanwhile, WavLM embeddings model high-level context-based and speaker-specific information. We employed Random Forest (RF), XGBoost (XGB), and Multilayer Perceptron (MLP) classifiers to assess the fused feature sets. The hybrid model, using MLP, achieved an accuracy of 95.9% and an Equal Error Rate (EER) of 0.063 on the ASVspoof2019 dataset. It also reached an accuracy of 97.4% with an EER of 0.029 on the Release-In-The-Wild dataset. These results show that the hybrid system is lightweight, interpretable, and very effective for practical deepfake audio detection.
Urban intersections are responsible for nearly 40
Efficient recognition of suspicious activity in surveillance videos requires reliable temporal modeling under strict computational constraints. Although uniform binarization significantly reduces computational complexity, it often disrupts temporal feature continuity, which is critical for video-based anomaly detection. This letter proposes a temporal-aware partially binarized vision transformer (TAPBViT) that introduces a signal-driven precision allocation strategy for spatiotemporal transformers. The proposed method models intermediate representations as temporal feature signals and assigns computational precision according to temporal sensitivity, preserving full precision temporal attention layers while binarizing spatial projection and feed-forward components that exhibit lower temporal distortion. This selective quantization mitigates quantization-induced temporal degradation while enabling aggressive computational reduction. Experiments on the UCF-Crime and UCSD Ped2 datasets demonstrate that TAPBViT achieves accuracies of 92.67% and 98.72%, respectively, while reducing binary operations by up to 10 & times; and supporting real-time inference on edge hardware. These results establish temporal signal-guided precision allocation as an effective principle for efficient transformer-based video surveillance.
As digital marketing continues to dominate global outreach strategies, worldwide spending on digital advertising is projected to surpass 785 billion by 2026. This surge is driven by real-time data analytics, behavioral targeting, and AI-powered personalization, all of which rely heavily on the large-scale collection and processing of consumer data. However, growing regulatory pressure from data privacy laws such as GDPR and CCPA, along with increasing public scrutiny, has elevated the demand for privacy-preserving and decentralized analytics frameworks. This study introduces an adaptive federated learning FL framework tailored for ordinal classification in digital marketing environments. The proposed system integrates two ordinal classifiers CORAL and CLM with a novel adaptive aggregation strategy that assigns dynamic weights to clients based on their contribution relevance, measured via feature importance. The experimental setup simulates a realistic collaborative marketing scenario involving five federated clients, each handling either real-world as Google Merchandise Store, UK Online Retail or synthetic as influencer and email campaign datasets. Through extensive experimentation, the framework demonstrates strong generalization across synthetic and real datasets, achieving classification accuracies up to 93.9
The complexity of traffic flow patterns significant challenges in predicting traffic green signal timings using conventional methods. Most of conventional methods relied on vehicle counts and speeds. These methods often did not consider crucial factors such as Spatial Occupancy, long-term dependencies, and the non-linear relationships. Recent advancements in Convolutional Neural Networks (CNNs) have enabled better capturing of patterns in traffic data. These advancements are essential for effectively predicting vehicle Green Signal Time by considering accurate detection and tracking, Spatial Occupancy calculation, long-term dependencies, and non-linear relationships in traffic data. The PVD-GSTPS framework has been proposed as an innovative solution for predicting vehicle Green Signal Time with the help of advanced CNN. This framework leverages the capabilities of two fine-tuned object detection models YOLO v8 and Faster R-CNN for precise vehicle detection, while a Byte Sort Tracker monitors the trajectories of detected vehicles. Additionally, a vehicle counting module assesses the number of vehicles in specified areas, and a size assignment process estimates Green Signal Time based on Spatial Occupancy calculations. This study is limited by the fixed duration of the QMUL video dataset utilized. This restricts data availability and complicates the establishment of strong correlations between Green Signal Time and Spatial Occupancy. To mitigate this issue, we utilized a Generative Adversarial Network (GAN) to generate realistic synthetic data. Long Short-Term Memory (LSTM) networks and polynomial regression techniques are utilized to capture the relationships within this dataset. In this study, we used the QMUL dataset to validate our hypothesis. The results demonstrate that our PVD-GSTPS framework significantly outperforms Enhanced YOLO v8, Original YOLO v8, and Faster R-CNN.
Image captioning generates text in natural language to describe a given image. Recent advances in various object detection with attention mechanisms pushed for exploring multiple image captioning methodologies to create more meaningful and accurate captioning models. Although existing pipelines do well in describing the image, there has not been enough emphasis on relationship modeling between image features. Relationship modeling is crucial for establishing context between various objects in an image. We have proposed a novel Advanced Context-Aware Object Relational Model (ACAORM) that not only improves relationship modeling between image features but also better sentence generation due to Transformer architecture. ACAORM builds relation-aware visual representations for image description and builds captions with prior attention to relevant Regions of Interest (RoI). We tested the proposed methodology with three widely used datasets, Flickr8K, Flickr30K, and MS-COCO. The results show that it beats numerous cutting-edge techniques. ACAORM scores 0.3526, 0.4439, and 0.8813 in BLEU-4 on Flickr8k, Flickr30k and MS-COCO, respectively, outperforming popular cutting-edge models such as VitaCap on Flickr8k and AGF on Flickr30k and MS-COCO by 10.88
Real-time traffic management systems are needed to manage urbanization’s impact on traffic conditions. Traffic dynamics in cities are complex, and traditional signal timing methods and simple vehicle detection cannot handle them. The paper describes a method for improving urban traffic studies using Real-time Dense Analysis and Management using Visual Graph Networks (RDAMVGN) that utilizes deep learning techniques along with Visual Graph Networks based on visualizations. This study aims to develop a robust, dynamic, and accurate approach to traffic density analysis, vehicle classification, and dynamic signal control in order to achieve high accuracy in traffic flow analysis. The proposed RDAMVGN framework incorporates both a LACF-YOLO detection model and a Faster Region-Based Convolutional Neural Network (Faster RCNN) detection model for high-speed and high-accuracy vehicle identification. This framework applies transfer learning to adapt pre-trained features to new traffic environments. This enhances vehicle classification in complex urban scenes and improves the model’s ability to distinguish vehicles from non-vehicle objects. Traffic flow optimization is achieved by using Mask RCNN and LSTM. The comparative analysis includes Fine-Tuned YOLOv8 with Fine-Tuned Faster RCNN, Fine-Tuned YOLOv3 with Fine-Tuned Faster RCNN, YOLOv5 with Faster RCNN, YOLOv3 with Faster RCNN, standalone Faster RCNN, YOLOv5, YOLOv3, Reinforcement Learning (RL), and Deep Reinforcement Learning (DRL) models, evaluate across precision, recall, accuracy, under pre-emption and non-pre-emption scenarios. The RDAMVGN-based detection component exhibits demonstrably superior performance across all evaluation metrics in real-time traffic management systems. It achieves a high precision of 97.75 %, indicating accurate vehicle identification with minimal false positives. Its recall rate of 96.17 % reflects strong detection capability, minimizing missed vehicles. The overall accuracy stands at 98.48 %, indicating robust classification and localization. In pre-emption scenario, the model maintains its lead with pre-emption precision of 94.79 %, pre-emption recall of 98.66 %, and pre-emption accuracy of 97.43 %, showcasing its reliability in real-time traffic prioritization scenarios.
Most Smartphone users prefer their phones to read news via various social platforms on the Internet. The news site publishes news and provides the source of identity verification. Humans are ineffective in distinguishing between true and false facts and make false news a threat to the logical truth, undermining the credibility of democracy, journalism, and government agencies. With the development of new technologies, finding a strategy to reduce the spread of false news or rumors that could harm society in any way is critical. Online clients are usually vulnerable to attack, and all content they run on web-based network media is generally considered reliable. Therefore, mechanized identification of fake news is essential to maintain a significant online media and informal organization. Believing in rumors and pretending to be news is harmful to society. A lot of effort is required to stop the rumors and focus on correct, verifiable news pieces, especially in developing countries. As a result, we offer a model for detecting fake news, a computational style analysis based on natural language processing that can effectively recognize text retrieved from social media Fake News using hybrid Hybrid model-based deep learning techniques.
Ensuring fairness in automated decision-making is a critical challenge, especially in organizational contexts like recruitment, performance evaluation, and promotion. As machine learning (ML) and artificial intelligence (AI) increasingly influence such decisions, promoting responsible AI that minimizes bias while preserving data privacy has become essential. However, existing fairness-aware models are often centralized or ill-equipped to handle non-IID data, limiting their real-world applicability. This study introduces a novel federated learning framework, Fairness-Weighted Federated Aggregation (FWFA), which integrates fairness-aware weighting into the model aggregation process. Each client’s contribution is scaled using a fairness score computed from key metrics Demographic Parity (DP), Statistical Parity Difference (SPD), and Disparate Impact Ratio (DIR). A synthetically generated dataset simulating diverse employee profiles across five professional domains was used to replicate real-world heterogeneity and imbalance. Across 20 communication rounds, FWFA achieved a DP of 0.91, an SPD of 0.06, and a DIR of 0.91, outperforming baseline methods WA+FL and SMOTE+FL, while maintaining an accuracy of 0.84. Additionally, a dynamic weighting mechanism was simulated by varying fairness thresholds to explore adaptive aggregation behavior, revealing a controllable trade-off between fairness and model performance. To further strengthen privacy guarantees, differential privacy was integrated into the FWFA framework, resulting in minimal performance degradation while retaining key fairness properties. These findings reinforce FWFA’s role as a robust, privacy-preserving solution for fair collaborative decision-making in federated environments, supporting the broader vision of ethical and trustworthy AI in real-world systems.
Federated learning (FL) presents a promising paradigm for decentralized machine learning, particularly well-suited for data-sensitive cyber-physical systems (CPS) where privacy preservation and low-latency inference are paramount. However, FL deployments at the edge face acute challenges in fault tolerance and communication efficiency due to the inherent unpredictability and heterogeneity of edge environments. This research addresses these issues through the development of a hierarchical federated learning framework that strategically combines clustered edge servers, the transfer-learning capabilities of pre-trained EfficientNet-B0 and TabNet models, and an enhanced aggregation approach based on FedProx to deliver superior robustness and efficiency. The proposed architecture introduces a multi-cluster aggregation mechanism, where edge servers within clusters coordinate local model updates before participating in higher-level aggregation, reducing inter-device communication overhead and overall latency. A distinctive contribution of this work is the explicit modeling and evaluation of system reliability by defining a "paralysis ratio" to quantify failed or unresponsive edge servers, and conducting real-time simulation experiments with incremental data delivery. Through this, the framework's resilience is demonstrated, maintaining model accuracy above 93% even under up to 40% server paralysis. In comparative benchmarks against the established Multi-Cluster Hierarchical FL (MCHFL) baseline, the framework achieves a 4-6% increase in classification accuracy and a 25% reduction in average communication latency. Beyond quantitative improvements, the study analyzes hyperparameter trade-offs, notably the impact of the FedProx constraint (mu), and highlights the practical value of real-time transfer learning EfficientNet-B0 for images and TabNet for complex time-series accelerating convergence under non-IID conditions. Collectively, these innovations advance FL by delivering a fault-tolerant, latency-aware, and transferable aggregation strategy validated through real-time simulation, significantly broadening the applicability of FL in complex, real-world edge and CPS deployments.