Content-Based Image Retrieval (CBIR) systems aim to retrieve images from databases based on their visual features, circumventing the limitations of traditional metadata-based approaches. This study introduces a novel approach that integrates features from Modified MobileNet and EfficientNetB0 to generate a comprehensive feature vector. By concatenating features from the intermediate and last layers of these models, the proposed approach captures both high-level semantic information and low-level details. The features extracted from Modified EfficientNetB0 have been optimized by using selective intermediate layer feature extraction. The selection of these intermediate layers has been performed by extensive experimentation to enhance feature-based representation and retrieval performance. The efficacy of the proposed approach is evaluated using similarity measures such as Euclidean distance and unsupervised nearest neighbors techniques on benchmark datasets, including Corel-1K and Caltech-101. The results demonstrate that the proposed method achieves 100 https://github.com/sunita2207/Optimization_of_intermediate_layers .
ABSTRACT Explainable artificial intelligence (XAI) is emerging as a critical enabler in cancer care, where high‐stakes decisions demand transparency, trust and regulatory accountability. This domain‐wise review systematically synthesises the adoption of XAI across major cancer types, classifying them into highly explored, moderately explored and underexplored categories. Various leading interpretability methods, including SHapley Additive exPlanations (SHAP), Local Interpretable Model‐agnostic Explanations (LIME), Gradient‐weighted Class Activation Mapping (Grad‐CAM) and counterfactual reasoning are critically examined, with their applications evaluated across imaging, genomics and multimodal cancer workflows. Certain persistent challenges such as data scarcity, methodological inconsistency, algorithmic bias and clinical validation gaps are highlighted, with particular focus on underrepresented cancers such as prostate, thyroid and pancreatic malignancies. The robustness and reproducibility of widely adopted XAI tools are evaluated, alongside an analysis of regulatory imperatives under emerging frameworks such as the European Union Artificial Intelligence Act (EU AI Act) and guidance from the United States Food and Drug Administration (FDA). In addition, this review outlines strategic directions through the integration of technical, clinical and ethical dimensions for the development of transparent, reliable and clinically relevant AI systems in cancer.
Background: Cervical cancer is among the most prevalent malignancies in women worldwide, and early detection of epigenetic alterations such as Deoxyribose Nucleic Acid (DNA) methylation is of utmost significance for improving clinical results. This study introduces a novel deep learning-based framework for predicting DNA methylation in cervical cancer, utilizing a UNet architecture integrated with an innovative one-hot character encoding technique. Methods: Two encoding strategies, monomer and dimer, were systematically evaluated for their ability to capture discriminative features from DNA sequences. Experiments were conducted on Cytosine–Guanine (CG) sites using varying sequence window sizes of 100 bp, 200 bp, and 300 bp, and sample sizes of 5000, 10,000, and 20,000. Model validation was performed on promoter regions of five cervical cancer-associated genes: miR-100, miR-138, miR-484, hTERT, and ERVH48-1. Results: The dimer encoding strategy, combined with a 300-base pair window and 5000 CG sites, emerged as the optimal configuration. The proposed framework demonstrated better predictive performance, with an accuracy of 91.60%, sensitivity of 96.71%, specificity of 87.32%, and an Area Under the Receiver Operating Characteristic (AUROC) score of 96.53, significantly outperforming benchmark deep learning models, including Convolutional Neural Networks and MobileNet. Validation on promoter regions further confirmed the robustness of the model, as it accurately identified 86.27% of methylated CG sites and maintained a strong AUROC of 83.99, demonstrating its precision–recall balance and practical relevance during validation in promoter-region genes. Conclusions: These findings establish the potential of the proposed UNet-based approach as a reliable and scalable tool for early detection of epigenetic modifications. Thus, the work contributes significantly to improving biomarker discovery and diagnostics in cervical cancer research.
Background: The purpose of the proposed study is to investigate the efficacy of UNet in predicting Deoxyribonucleic Acid methylation patterns in a cervical cancer cell line. The application of deep learning to analyse the factors affecting methylation in the context of cervical cancer has not yet been fully explored. Methods: A comprehensive performance evaluation has been conducted based on multiple window sizes of DNA sequences. For this purpose, three different parameter-analysis techniques, namely, autoencoders, Generative Adversarial Networks, and Multi-Head Attention Networks, were used. This work presents a novel framework for methylation prediction in promoter regions of various genes. Results and Conclusions: Experimental results have proved that attention networks in association with UNet achieved a significant accuracy level of 91.01% along with a sensitivity of 89.65%, specificity of around 92.35%, and an area under curve of 0.910 on ENCODE database. The proposed model outperformed three state-of-the-art models: Convolutional Neural Network, Transfer Learning, and Feed Forward Neural Network with K-Nearest Neighbour. Moreover, validation of the model in five gene promoters achieved an accuracy of 81.60% with an area under curve score of 0.814, a p-value of 3.62×10-19, and Cohen's Kappa value of 0.631. This novel approach has led to a better understanding of epigenetic variables and their implications in cervical cancer, offering potential insights into therapeutic strategies.
Underwater images play a crucial role in advancing marine research, enhancing surveillance and ensuring secure communication. However, underwater image security is a significant concern due to the unique transmission and distortion vulnerabilities present in the aquatic environment. Factors such as light scattering, absorption, turbidity and limited bandwidth can degrade image quality, distort visual features and increase vulnerability to cyber threats. These environmental challenges make it difficult for existing encryption methods to simultaneously ensure high security and preserve image fidelity under submerged conditions. In response this study presents UIGen, a novel and secure encryption method specifically designed for underwater images that addresses the inherent security challenges of the aquatic environment while ensuring the confidentiality and integrity of the images. It employs Non-Subsampled Shearlet Transform (NSST) technique to extract multi-scale features from underwater images, enabling effective representation of both coarse and fine details. A dynamic key generation approach based on a hyper-chaotic parameter is used to increase key sensitivity and strengthen the overall encryption security. Finally, a Genetic Algorithm (GA) is applied to optimize the encryption parameters, enhancing the robustness of the system and ensuring its adaptability to varying underwater conditions. Experimental results demonstrate that UIGen maintains robust structural integrity and strong entropy while resisting underwater-specific attacks and enabling high-fidelity image encryption and storage.
Traffic congestion has become a major issue in growing Smart Cities. Bike-sharing systems offer an eco-friendly solution to reduce congestion, but predicting bike demand accurately is still a difficult task because it depends on many changing factors like weather, season, and time of day. Many existing models either use all features without checking their importance or fail to adapt to real-time data, leading to less reliable predictions. To overcome these challenges, this study introduces a hybrid Grey Wolf Optimization-based Incremental Extreme Learning Machine (GWO–IELM) model for predicting bike-sharing demand. In this framework, GWO selects the most important features to reduce unnecessary data and computation, while IELM makes fast and accurate predictions by learning incrementally. The model is evaluated on the Kaggle London bike-sharing dataset, which includes over 17,000 hly records with meteorological, temporal, and holiday-related features. The experimental results show that GWO-IELM consistently outperforms standard models including Linear Regression, Support Vector Regressor, AdaBoost Regressor, and Bagging Regressor with an R ^2 of 0.9897, MAE of 0.0859, RMSE of 0.1547, and RMSLE of 0.0431. The robustness is further validated using the Diebold-Mariano test. Further, the interpretability of model is supported through SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME), providing insights into the contribution of selected features. These results demonstrate the potential of integrating optimization, lightweight learning, and explainability for improving demand prediction in bike-sharing systems and supporting smarter urban mobility strategies.
Summary The sensor nodes, which are available in the wireless sensor networks (WSN), are equipped with sensing abilities, and communication. Several domains require the sensor nodes to be arranged in aggressive surroundings, in which observing malicious activities within the sensor network. Therefore, the present research proposes the border‐hunting optimization‐based deep CNN (BHO‐DCNN) for the mobile agent (MA)‐based intrusion detection in WSN. The importance of the research relies on the BHO‐DCNN model for identifying the intrusion available in the network is established by integrating the optimization with its features through a deep classifier for detection in a precise manner. The algorithm follows the communal hierarchy, surrounding, group hunting, and prey attacking, which provides an enhanced rate of convergence in the method of detection. The analysis is achieved through the database IDS 2018 Intrusion CSVs depending on the performance like delay, alive nodes, normalized energy, as well as throughput. The obtained number of alive nodes through the developed BHO‐DCNN algorithm is 45, end‐to‐end delay is 0.2572 ms, normalized energy is 0.1622 J, as well as throughput, is 0.3125% for nodes 50 at the populate rate of 100, respectively.
Sensor nodes can be deployed in harsh or hostile environments in many applications, making these nodes more prone to failure. The illegal movement monitoring within the sensor networks is a most challenging problem. The mobile malicious nodes are preferred by the attacker to maximize his impact. For a dynamic environment, a promising technology of sensor networks is expected to Intrusion detection. Multi-mobile agents utilize many approaches, after verification that collects data from sensor nodes. However, these approaches are inefficient to verify all the sensor nodes (SNs) of the network, due to its high delay, energy consumption, and mobility. The proposed Dunnock Ibis optimization LSTM model (DIO opt LSTM) solves this problem. Here, the sensor nodes are grouped into clusters; hence, mobile agent performs verification only the cluster heads instead of verifying all the SNs. The proposed DIO optimization combines the unique behavior of Egret Swam and Ibis optimization algorithm which efficiently tunes the LSTM classifier, resulting in the model providing better convergence. The simulation results show the proposed system shows a better result than the existing system by utilizing the database IDS 2018 Intrusion CSVs, the analysis is done based on performance metrics such as End-end-delay (ED), normalized energy (NE), and throughput. At 200 nodes and 1500 rounds, the DIO opt LSTM method has efficiently performed 146 numbers of alive nodes, 0.46ms of delay, 0.15J of normalized energy, and 0.89 bps of throughput.
This study investigates the ability to predict DNA methylation patterns in cervical cancer cells using decision-tree-based ensemble approaches and neural network-based models. The research findings suggest that a model based on random forest achieves a significant prediction accuracy of 91.35%. This projection was derived from comprehensive experimentation and a meticulous performance evaluation of the random forest model, employing a range of measures including Accuracy, Sensitivity, Specificity, Matthews Correlation Coefficient, F1-score, Recall, and Precision. The results indicate that the random forest model exhibits superior performance compared to other tree-based models such as the Simple Decision Tree and XGBoost, as well as neural network-based models including Convolutional Neural Networks, Feed Forward Networks, and Wavelet Neural Networks. The findings indicate that using random forest-based techniques has great potential for future study and might be highly valuable in clinical applications, especially in improving diagnostic and treatment strategies based on epigenetic profiles.
Analysing patterns/trends and associations from het-erogeneous data coming at varied speeds and formats require data structures which can handle large and dynamic data efficiently. Bloom Filter (BF), a probabilistic data structure, can be considered as one of the alternatives as it is space efficient and handles membership query effectively. To accommodate the dynamic data sets in the BF when ever the Fill Ratio (FR) exceeds the defined limits, two major approaches have been discussed in the literature: increasing the number of BFs as the data increases and refreshing the existing BF by deleting the staled data. In this paper, a novel approach motivated from Double probing named (DH-FRP) has been proposed to compute FR using augmented array to sample bit locations from parent BF. The augmented array keeps track of the sampled bit values at any given time and stores at contiguous locations for effcient bit counting and FR calculation. Markov Bound has been used to determine whether to recompute the step size or not. Since the rate of dynamic data coming from various sources may be uniform or non uniform, both cases have been considered and FR has been calculated for standard BF and Counting BF. It has been experimentally proved that the proposed scheme outperforms the given state-of-art techniques in terms of FR calculation time and accuracy.
Content-Based Image Retrieval (CBIR) is a method for retrieving images based on their content rather than relying on textual descriptions or tags. Over the last decade, Deep Convolutional Neural Networks (D-CNN) based architectures have gained popularity in image retrieval. One major issue with D-CNN is that they are complex, heavy networks with high dimensional features. To overcome this bottleneck, a separable convolutional neural networks based framework has been proposed, which will reduce the complexity of the network. The proposed framework minimizes the length of the final feature vector by extracting the detailed features from the intermediate layers. This step facilitates the avoidance of abstraction of features obtained from the last layer only. The proposed technique has achieved recall and accuracy of the order of 1, indicating a high degree of relevant retrieval. For indexing and similarity matching, the approximate Nearest Neighbour Search (ANNOY) technique has been applied and optimal retrieval has been achieved when compared with the existing techniques, indicating the better performance of the proposed technique. Experimental results indicate that the proposed framework performs exceptionally well compared to the existing state-of-the-art techniques and can be easily employed in various image-related applications.
Majority of the learning algorithms used for the training of feedforward neural networks (FNNs), such as backpropagation (BP), conjugate gradient method, etc. rely on the traditional gradient method. Such algorithms have a few drawbacks, including slow convergence, sensitivity to noisy data, local minimum problem, etc. One of the alternatives to overcome such issues is Extreme Learning Machine (ELM), which requires less training time, ensures global optimum and enhanced generalization in neural networks. ELM has a single hidden layer, which poses memory constraints in some problem domains. An extension to ELM, Multilayer ELM (ML-ELM) performs unsupervised learning by utilizing ELM autoencoders and eliminates the need of parameter tuning, enabling better representation learning as it consists of multiple layers. This paper provides a thorough review of ML-ELM architecture development and its variants and applications. The state-of-the-art comparative analysis between ML-ELM and other machine and deep learning classifiers demonstrate the efficacy of ML-ELM in the niche domains of Computer Science which further justifies its competency and effectiveness.
Abstract The latest technology trend is the conflation of mobile agent technology for high-level inference and surveillance in wireless sensor networks. Bridging the two technology contributes to the emerging areas and acquiring foremost importance in communication. The sensor nodes shift from fixed to a dynamic environment that changes rapidly becomes prone to malicious attacks. This work considers the attacks which cause the communication failure and raises the alarm for appropriate action. A scheme has been defined to collect sufficient data to identify the abnormal behaviour on sensor nodes from Network Simulator (NS-2.35) on the defined features with mobile agent based intrusion detection system (IDS) using SPIN protocol is one of the most useful data-centric routing protocols in WSNs. This work describes a novel approach for classifying attacks based on consumed energy onto different parameters. Furthermore, on the basis of attacks type, a performance matrix is computed by standard classifiers with machine learning models. An ensemble model is proposed to improve the performance of the ensemble classifiers, which yields an average accuracy of 95%. Further, K Fold cross validation has been performed to check the consistency of the proposed ensemble model.
Named Data Networking (NDN) has gained importance in today’s era due to a paradigm shift in the Internet usage pattern which revolves around the content rather than the respective host addresses. Three important data structures in NDN are Content Store (CS), Pending Interest Table (PIT) and Forwarding Information Base (FIB). The search time of PIT is quite high since its size grows with the addition of new content names, and the Interest packets which are not served by CS are searched in millions of existing entries in the PIT. Hence, lookup time can be improved if, instead of checking all the available entries, initial scanning is done to determine whether the required content name exists in the PIT or not. In this paper, we propose a Stable Bloom Filter (SBF) based PIT called S-PIT, to minimize the PIT search time by identifying the existence of query content through SBF. The various experiments performed show that S-PIT outperforms existing data structures in terms of memory consumption, content insertion time, average search time and false positive rate.
With each passing day, the number of smart vehicles is increasing manifold, hence, automatic/automated parking lot detection is gaining a lot of importance among Smart City applications. A robust approach is desired to identify parking spaces effectively and efficiently. This work presents a deep learning classifier based on convolutional neural network (CNN) and extreme learning machine (ELM), i.e., CNN-ELM to classify the parking space as vacant or occupied. CNN is well-known for efficient image classification, but its training time is highly influenced by backpropagation of errors in the fully connected layer. Thus, ELM is plugged into CNN to replace the fully connected layer and perform classification whereas, feature extraction is performed using CNN. The performance of CNN-ELM is validated on the publicly available PKLot dataset, which contains approximately 700,000 images categorized into sunny, overcast, and rainy weather conditions. The experimental results indicate that CNN-ELM approach outperforms other hybrid CNN models using different classifiers such as support vector machine, Xgboost, and Extra Trees in terms of sensitivity, specificity, and accuracy. The comparison of results with other state-of-the-art approaches based on accuracy and Area under the curve (AUC) score further justifies the effectiveness of the proposed approach in real-time parking space detection.
Essential emergency services can be provided, and many lives can be saved if the severity of the road accident is analyzed well in time. Several works have been proposed to ascertain accident severity in intelligent transportation system based on traditional machine learning (ML) approaches such as Logistic Regression, Support Vector Machines (SVM), Random Forests, etc. The motive of this research is to propose an efficient and reliable approach for classifying the severity of road accidents through combined techniques of the feature space of extreme learning machine (ELM) and SVM, named as E-SVM, by making the best use of their characteristics. ELM performs feature mapping of input data, and further, the radial basis function kernel is utilized for training the SVM model, which performs the classification process. The Extra-Trees classifier used for feature selection leads to a reduced dataset of significant features which contributes towards efficient accident severity classification. The experimental results show that the proposed approach is better as compared to the other state-of-the-art ML classifiers both in terms of computational time and system performance, hence justifying its usage in real-life applications.
Concept drift refers to the change in data distributions and evolving relationships between input and output variables with the passage of time. To analyze such variations in learning environments and generate models which can accommodate changing performance of predictive systems is one of the challenging machine learning applications. In general, the majority of the existing schemes consider one of the specific drift types: gradual, abrupt, recurring, or mixed, with traditional voting setup. In this work, we propose a novel data stream framework, dynamically adaptive and diverse dual ensemble (DA‐DDE) which responds to multiple drift types in the incoming data streams by combining online and block‐based ensemble techniques. In the proposed scheme, a dual diversified ensemble‐based system is constructed with the combination of active and passive ensembles, updated over a diverse set of resampled input space. The adaptive weight setting method is proposed in this work which utilizes the overall performance of learners on historic as well as recent concepts of distributions. Further a dual voting system has been used for hypothesis generation by considering dynamic adaptive credibility of ensembles in real time. Comparative analysis with 14 state‐of‐the‐art algorithms on 24 artificial and 11 real datasets shows that DA‐DDE is highly effective in handling various drift types.
Internet of Vehicles (IoV) has escalated the movement of big data across moving vehicles which create a huge burden on the network infrastructure. In IoV environment, effective handling of streaming data has to face various challenges like; traffic monitoring, flow management, re-configuration and security. Software-defined networks (SDN) provides improved flexibility, and centralized control of the network to overcome (almost) the above-mentioned challenges. However, it can lead to an easy target (node or controller) for malicious agents. So, to detect the anomalous behaviour of the nodes in the IoV environment, a hybrid approach using probabilistic data structures is proposed which works in the following phases. In phase I, a traffic monitoring scheme using Count-Min-Sketch is designed to identify the suspicious nodes. In phase II, to detect an anomaly, a Bloom filter-based control scheme is used for signature verification of suspicious nodes. In phase III, a Quotient filter is used for fast and efficient storage of malicious nodes. In phase IV, to detect the super points (malicious hosts that are connected to a large number of destinations), a Hyperloglog counter is used to measure the cardinality of each flow passing through the switches. The proposed scheme has been evaluated in a simulated environment. The results obtained depict that the proposed scheme is faster, accurate, and efficient concerning detection ratio and false-positive ratio.
In this paper, traditional and meta-heuristic approaches for optimizing deep neural networks (DNN) have been surveyed, and a genetic algorithm (GA)-based approach involving two optimization phases for hyper-parameter discovery and optimal data subset determination has been proposed. The first phase aims to quickly select an optimal combination of the network hyper-parameters to design a DNN. Compared to the traditional grid-search-based method, the optimal parameters have been computed 6.5 times faster for recurrent neural network (RNN) and 8 times faster for convolutional neural network (CNN). The proposed approach is capable of tuning multiple hyper-parameters simultaneously. The second phase finds an appropriate subset of the training data for near-optimal prediction performance, providing an additional speedup of 75.86% for RNN and 41.12% for CNN over the first phase.
In the present era many real world applications are built from streaming data. The distribution of underlying data in such streams tend to change with course of time called as concept drift. A lot of algorithms are proposed in machine learning and data mining domain which are used to handle this drifting scenario. Ensembles form an integral component of such algorithms. In this paper, randomization is added using online bagging to the existing four drift handling approaches and its effect is analysed over multiple patterns of concept drift such as gradual, abrupt, recurring etc. Experimental work has been conducted over artificially generated data streams and real datasets to validate the impact of bagging in learning process.