The progressive dissemination of the Internet of Things (IoT) and Wireless Sensor Networks (WSNs) has ushered in a new era of connectivity, with vast applications spanning from medicare to smart city infrastructure. However, this expansion has been paralleled by a corresponding increase in the sophistication and variety of cyber threats targeting these networks. Traditional cyber security measures, designed for a less dynamic threat landscape, are proving increasingly insufficient in protecting against the innovative and varied attack methods now in commonplace. This study introduces an innovative application of Generative Adversarial Networks (GANs) to address this challenge, presenting a novel framework for the simulation and mitigation of advanced network attacks, particularly focusing on Distributed Denial of Service (DDoS) and spoofing attacks which pose significant threats in IoT environments. Generative Adversarial Networks (GANs), comprising two neural networks-the generator and the discriminator compete in a game-theoretic scenario, facilitating a deep understanding of attack patterns through the generation of realistic, synthetic cyber-attack scenarios. This research exploits GANs to bridge the gap between the static nature of traditional security protocols and the dynamic, evolving landscape of cyber threats. By training on a comprehensive dataset of known attacks and normal network activities, our proposed model, the Dynamic Adaptive Threat Simulation GAN (DATS-GAN), is capable of producing varied and realistic attack scenarios. These simulations serve a dual purpose: they not only enhance the detection capabilities and responsiveness of current security systems but also provide a basis for the development of new, adaptive security mechanisms capable of dynamically responding to the ever-changing cyber threat landscape. The effectiveness of DATS-GAN is demonstrated through extensive empirical analysis, highlighting significant improvements in the detection precision and reaction times of security frameworks within WSNs. Moreover, the generated synthetic attack scenarios provide a valuable resource for training machine learning models, leading to the advancement of adaptive security solutions that maintain a high readiness level against emerging cyber threats. The outcomes of this research hold substantial promise for the cyber security domain, showcasing the potential of GANs to revolutionize network defenses against sophisticated cyber threats in IoT and WSN environments.
Background/Objectives: This study presents a comparative analysis of the multistage diagnosis of Alzheimer’s disease (AD), including mild cognitive impairment (MCI), utilizing two distinct types of biomarkers: blood gene expression and clinical biomarker samples. Both of these samples, obtained from participants in the Alzheimer’s Disease Neuroimaging Initiative (ADNI), were independently analyzed utilizing machine learning (ML)-based multiclassifiers. This study applied novel machine learning-based data augmentation techniques to gene expression profile data that are high-dimensional, low-sample-size (HDLSS) and inherently highly imbalanced. The investigation obtained the highest multiclassification performance to date in the multistage diagnosis of Alzheimer’s disease utilizing the blood gene expression profiles of Alzheimer’s Disease Neuroimaging Initiative (ADNI) participants. Based on the performance results obtained, and other factors such as early prediction capabilities, this study compares the efficacies of the two types of biomarkers for multistage diagnosis. This study presents the sole investigation in which multiclassification-based AD stage diagnosis was conducted utilizing blood gene expression data. We obtained the best multiclassification result in both modalities of the ADNI data in terms of F1-score and were able to identify new genetic biomarkers. Methods: The combination of the XGBoost and SFBS (Sequential Floating Backward Selection) methods was used to select the features. We were able to select the 95 most effective gene probe sets out of 49,386. For the clinical study data, eight of the most effective biomarkers were selected using SFBS. A deep learning (DL) classifier was used to identify the stages—cognitive normal (CN), mild cognitive impairment (MCI), and Alzheimer’s disease (AD)/dementia. DL, support vector machine (SVM), gradient boosting (GB), and random forest (RF) classifiers were used for the AD stage detection from gene expression profile data. Because of the high data imbalance in genomic data, borderline oversampling/data augmentation was applied in the model training and original samples for validation. Results: Utilizing clinical data, the highest ROC AUC scores attained were 0.989, 0.927, and 0.907 for the identification of the CN, MCI, and dementia stages, respectively. The highest F1 scores achieved were 0.971, 0.939, and 0.886. Employing gene expression data, we obtained ROC AUC scores of 0.763, 0.761, and 0.706 for the CN, MCI, and dementia stages, respectively, and F1 scores of 0.71, 0.77, and 0.53 for CN, MCI, and dementia, respectively. Conclusions: This represents the best outcome to date for AD stage diagnosis from ADNI blood gene expression profile data utilizing multiclassification techniques. The results indicated that our multiclassification model effectively manages the imbalanced data of a high-dimension, low-sample-size (HDLSS) nature to identify samples of the minority class. MAPK14, PLG, FZD2, FXYD6, and TEP1 are among the novel genes identified as being associated with AD risk.
INTRODUCTION: This study focusses on diagnosis of stages of AD (Alzheimer’s disease) including MCI (Mild Cognitive Impairment) from two data modalities - gene expression and clinical data of ADNI (Alzheimer’s Disease Neuroimaging Initiative ) participants using multiclassification. The gene expression dataset is highly imbalanced and of HDLSS (high-dimensional and low-sample-size) characteristics. This is the only study where multiclassification based AD stage diagnosis is done to identify multiple stages of Alzheimer. We are able to achieve the best multiclassification result in both the modalities and identify new genetic biomarkers.METHODS: Combination of XGBoost and SFBS (“Sequential Floating Backward Selection”) methods is used to select features. We are able to select the most effective 95 gene probsets out of 49,386. For clinical study data, 8 most effective biomarkers could be selected using SFBS. For both genomic and clinical data, DL (‘Deep Learning’) classifier is used to identify stages - CN (Cognitive Normal), MCI (Mild Cognitive Impairment), AD (Alzheimer’s Disease / Dementia). Because of high data imbalance in genomic data, border line oversampling is used for model training and original data for validation.RESULT & DISCUSSION: With clinical data, we achieved ‘ROC AUC’ scores 0.97, 0.95, 0.94 for CN, MCI, Dementia stage respectively . We achieve ‘ROC AUC’ scores 0.75, 0.74, 0.70 for CN, MCI, Dementia stage respectively and 0.67 for both micro average F1 scores and micro weighted F1 score. This is the best result so far for AD stage diagnosis from gene expression profile data through multiclassification with ADNI data. Results reflect that our multiclassification model can efficiently handle the imbalanced data of HDLSS nature to identify samples of minority class. MAPK14, ZNF835, MID1, HLA-DQA1, TEP1 are some of the new genes found to be associated with AD risk. DRAXIN, HSPA12B, USP47 etc. are found to be AD preventive or suppressor.
INTRODUCTION: This study focusses on diagnosis of stages of AD (Alzheimer’s disease) including MCI (Mild Cognitive Impairment) from two data modalities - gene expression and clinical data of ADNI (Alzheimer’s Disease Neuroimaging Initiative ) participants using multiclassification. The gene expression dataset is highly imbalanced and of HDLSS (high-dimensional and low-sample-size) characteristics. This is the only study where multiclassification based AD stage diagnosis is done to identify multiple stages of Alzheimer. We are able to achieve the best multiclassification result in both the modalities and identify new genetic biomarkers. METHODS: Combination of XGBoost and SFBS (“Sequential Floating Backward Selection”) methods is used to select features. We are able to select the most effective 95 gene probsets out of 49,386. For clinical study data, 8 most effective biomarkers could be selected using SFBS. For both genomic and clinical data, DL (‘Deep Learning’) classifier is used to identify stages - CN (Cognitive Normal), MCI (Mild Cognitive Impairment), AD (Alzheimer’s Disease / Dementia). Because of high data imbalance in genomic data, border line oversampling is used for model training and original data for validation. RESULT & DISCUSSION: With clinical data, we achieved ‘ROC AUC’ scores 0.97, 0.95, 0.94 for CN, MCI, Dementia stage respectively . We achieve ‘ROC AUC’ scores 0.75, 0.74, 0.70 for CN, MCI, Dementia stage respectively and 0.67 for both micro average F1 scores and micro weighted F1 score. This is the best result so far for AD stage diagnosis from gene expression profile data through multiclassification with ADNI data. Results reflect that our multiclassification model can efficiently handle the imbalanced data of HDLSS nature to identify samples of minority class. MAPK14, ZNF835, MID1, HLA-DQA1, TEP1 are some of the new genes found to be associated with AD risk. DRAXIN, HSPA12B, USP47 etc. are found to be AD preventive or suppressor.
This paper investigates the utilization of a physics simulation environment as the imagination of a robot, where it creates a replica of the detected terrain in a physics simulation environment in its memory, and “imagines” a simulated version of itself in that memory, performing actions and navigation on the terrain. The physics of the environment simulates the movement of robot parts and its interaction with the objects in the environment and the terrain, thus avoiding the need for explicitly programming many calculations.
Late-onset Alzheimer’s disease (LOAD) is a subtype of dementia that manifests after the age of 65. It is characterized by progressive impairments in cognitive functions, behavioral changes, and learning difficulties. Given the progressive nature of the disease, early diagnosis is crucial. Early-onset Alzheimer’s disease (EOAD) is solely attributable to genetic factors, whereas LOAD has multiple contributing factors. A complex pathway mechanism involving multiple factors contributes to LOAD progression. Employing a systems biology approach, our analysis encompassed the genetic, epigenetic, metabolic, and environmental factors that modulate the molecular networks and pathways. These factors affect the brain’s structural integrity, functional capacity, and connectivity, ultimately leading to the manifestation of the disease. This study has aggregated diverse biomarkers associated with factors capable of altering the molecular networks and pathways that influence brain structure, functionality, and connectivity. These biomarkers serve as potential early indicators for AD diagnosis and are designated as early biomarkers. The other biomarker datasets associated with the brain structure, functionality, connectivity, and related parameters of an individual are broadly categorized as clinical-stage biomarkers. This study has compiled research papers on Alzheimer’s disease (AD) diagnosis utilizing machine learning (ML) methodologies from both categories of biomarker data, including the applications of ML techniques for AD diagnosis. The broad objectives of our study are research gap identification, assessment of biomarker efficacy, and the most effective or prevalent ML technology used in AD diagnosis. This paper examines the predominant use of deep learning (DL) and convolutional neural networks (CNNs) in Alzheimer’s disease (AD) diagnosis utilizing various types of biomarker data. Furthermore, this study has addressed the potential scope of using generative AI and the Synthetic Minority Oversampling Technique (SMOTE) for data augmentation.
Late-onset Alzheimer’s disease (LOAD) is a dementia category disease that starts after the age of 65. Progressive problems in thinking, behaviour, learning etc. are its characteristics. Due to its progressive nature, early diagnosis is crucial. A complex pathway mechanism with different risk factors is involved in LOAD progression. A comprehensive study is done on factors causing LOAD, biomarkers used for diagnosis in clinical and preclinical stage of the disease, ML (Machine Learning) techniques to identify disease stage. We reviewed research papers on clinical and preclinical diagnosis of LOAD using ML techniques from biomarker data: MRI, PET image, MMSE, Genome / Gene expression, DNA / RNA binding sites with protein, Epigenetic expression /DNA, methylation, Speech and text. We have listed improvement scope after finding the gaps and scope of using Generative AI and SMOTE (Synthetic Minority Oversampling Technique) for data augmentation. At the end, we have elaborated current use of DL (Deep Learning) and CNN (‘Convolution Neural Network’) in Alzheimer’s disease (AD) diagnosis.
We present Limousine, a self-designing key-value storage engine, that can automatically morph to the near-optimal storage engine architecture shape given a workload, a cloud budget, and target performance. At its core, Limousine identifies the fundamental design principles of storage engines as combinations of learned and classical data structures that collaborate through algorithms for data storage and access. By unifying these principles over diverse hardware and three major cloud providers (AWS, GCP, and Azure), Limousine creates a massive design space of quindecillion (1048) storage engine designs the vast majority of which do not exist in literature or industry. Limousine contains a distribution-aware IO model to accurately evaluate any candidate design. Using these models, Limousine searches within the exhaustive design space to construct a navigable continuum of designs connected along a Pareto frontier of cloud cost and performance. If storage engines contain learned components, Limousine also introduces efficient lazy write algorithms to optimize the holistic read-write performance. Once the near-optimal design is decided for the given context, Limousine automatically materializes the corresponding design in Rust code. Using the YCSB benchmark, we demonstrate that storage engines automatically designed and generated by Limousine scale better by up to 3 orders of magnitude when compared with state-of-the-art industry-leading engines such as RocksDB, WiredTiger, FASTER, and Cosine, over diverse workloads, data sets, and cloud budgets.
Kidney tumors represent a significant medical challenge, characterized by their often-asymptomatic nature and the need for early detection to facilitate timely and effective intervention. Although neural networks have shown great promise in disease prediction, their computational demands have limited their practicality in clinical settings. This study introduces a novel methodology, the UNet-PWP architecture, tailored explicitly for kidney tumor segmentation, designed to optimize resource utilization and overcome computational complexity constraints. A key novelty in our approach is the application of adaptive partitioning, which deconstructs the intricate UNet architecture into smaller submodels. This partitioning strategy reduces computational requirements and enhances the model's efficiency in processing kidney tumor images. Additionally, we augment the UNet's depth by incorporating pre-trained weights, therefore significantly boosting its capacity to handle intricate and detailed segmentation tasks. Furthermore, we employ weight-pruning techniques to eliminate redundant zero-weighted parameters, further streamlining the UNet-PWP model without compromising its performance. To rigorously assess the effectiveness of our proposed UNet-PWP model, we conducted a comparative evaluation alongside the DeepLab V3+ model, both trained on the "KiTs 19, 21, and 23" kidney tumor dataset. Our results are optimistic, with the UNet-PWP model achieving an exceptional accuracy rate of 97.01% on both the training and test datasets, surpassing the DeepLab V3+ model in performance. Furthermore, to ensure our model's results are easily understandable and explainable. We included a fusion of the attention and Grad-CAM XAI methods. This approach provides valuable insights into the decision-making process of our model and the regions of interest that affect its predictions. In the medical field, this interpretability aspect is crucial for healthcare professionals to trust and comprehend the model's reasoning.
Chronic Kidney Disease (CKD) represents a considerable global health challenge, emphasizing the need for precise and prompt prediction of disease progression to enable early intervention and enhance patient outcomes. As per this study, we introduce an innovative fusion deep learning model that combines a Graph Neural Network (GNN) and a tabular data model for predicting CKD progression by capitalizing on the strengths of both graph-structured and tabular data representations. The GNN model processes graph-structured data, uncovering intricate relationships between patients and their medical conditions, while the tabular data model adeptly manages patient-specific features within a conventional data format. An extensive comparison of the fusion model, GNN model, tabular data model, and a baseline model was conducted utilizing various evaluation metrics, encompassing accuracy, precision, recall, and F1-score. The fusion model exhibited outstanding performance across all metrics, underlining its augmented capacity for predicting CKD progression. The GNN model's performance closely trailed the fusion model, accentuating the advantages of integrating graph-structured data into the prediction process. Hyperparameter optimization was performed using grid search, ensuring a fair comparison among the models. The fusion model displayed consistent performance across diverse data splits, demonstrating its adaptability to dataset variations and resilience against noise and outliers. In conclusion, the proposed fusion deep learning model, which amalgamates the capabilities of both the GNN model and the tabular data model, substantially surpasses the individual models and the baseline model in predicting CKD progression. This pioneering approach provides a more precise and dependable method for early detection and management of CKD, highlighting its potential to advance the domain of precision medicine and elevate patient care.
Autism spectrum disorder (ASD) is a complex and heterogeneous neurodevelopmental disorder. Machine learning and deep learning techniques have been playing an important role in automating the diagnosis of brain disorder, which is characterized by social deficits and repetitive behaviors. In this paper, we have proposed and implemented a machine learning model and convolution neural network (CNN) for classifying subjects with ASD. Data are from Autism Brain Imagining Data Exchange (ABIDE) repository by using phenotypic, sMRI, and fMRI data. For sMRI image dataset, the accuracy of the neural network is about 87
The objective of this study was to develop a system for chronic kidney disease (CKD) and to identify relevant prognostic features using a clinical dataset. Accurate classification and major risk factors in chronic kidney disease lead to better prognosis and assist nephrologists. Due to privacy and other factors, the data source is not balanced to trail any models. Therefore, it is difficult to achieve consistent accuracy with an imbalanced dataset, and there will be a variance in results with different machine learning models. In the proposed study, GAN's generated synthesised dataset, which is very close to the original dataset, is used. A hybrid synthesised dataset consists of the original dataset along with the synthesised data generated with the GAN model. The proposed model also includes the most important risk variables for CKD. The metrics used in the study include F1-score, accuracy, and from the plot, it shows that the TabNet with GAN's synthetic data is more consistent and more accurate than the traditional machine learning techniques with imbalanced dataset. The proposed model iterated for 150 times to get the variance, which is much less than in proposed techniques with hybrid preprocessed datasets. The proposed work significantly increased the classification accuracy of chronic kidney disease. These models and parameters show how important health status data is for predicting the risk of and development of kidney disease.
Background: Accurate semantic segmentation of kidney tumors in computed tomography (CT) images is difficult because tumors feature varied forms and occasionally, look alike. The KiTs19 challenge sets the groundwork for future advances in kidney tumor segmentation. Methods: We present weight pruning (WP)-UNet, a deep network model that is lightweight with a small scale; it involves few parameters with a quick assumption time and a low floating-point computational complexity. Results: We trained and evaluated the model with CT images from 210 patients. The findings implied the dominance of our method on the training Dice score (0.98) for the kidney tumor region. The proposed model only uses 1,297,441 parameters and 7.2e floating-point operations, three times lower than those for other network models. Conclusions: The results confirm that the proposed architecture is smaller than that of UNet, involves less computational complexity, and yields good accuracy, indicating its potential applicability in kidney tumor imaging.
This research was conducted with the goals of developing models that have an accurate classification of chronic kidney disease (CKD) and locating significant prognostic factors within a clinical dataset. In chronic kidney disease, accurate classification and identification of major risk factors contribute to improved prognoses and provide assistance to nephrologists. The data source is not balanced enough to serve as a benchmark for any machine learning or deep learning models due to privacy concerns and other factors. As a result, it is difficult to achieve consistent accuracy with an imbalanced dataset, and there will be a variance in results with various machine learning models due to the fact that there are so many different factors involved. A multiple-level ensemble learning system was utilised in the proposed research in order to classify chronic kidney disease using a hybrid synthesised dataset. A hybrid synthesised dataset is one that contains both the original dataset as well as the synthesised data that was produced using the ADASYN method. The most important risk factors for chronic kidney disease are accounted for in the model that was proposed. The F1-score and accuracy were two of the metrics that were utilised in this research. Furthermore, the plot demonstrates that the multilevel ensemble is superior to the conventional machine learning techniques in terms of consistency and accuracy. The variance was calculated using the proposed model by iterating for a total of 150 times, which is a significant reduction when compared to multi-level ensemble techniques using hybrid preprocessed datasets. The accuracy of classifying patients with chronic kidney disease was significantly improved by using a multi-level ensemble. These models and parameters highlight the significance of current health status information in estimating the likelihood of developing kidney disease as well as its progression.
We present Sparse Numerical Array-Based Range Filters (SNARF), a learned range filter that efficiently supports range queries for numerical data. SNARF creates a model of the data distribution to map the keys into a bit array which is stored in a compressed form. The model along with the compressed bit array which constitutes SNARF are used to answer membership queries. We evaluate SNARF on multiple synthetic and real-world datasets as a stand-alone filter and by integrating it into RocksDB. For range queries, SNARF provides up to 50x better false positive rate than state-of-the-art range filters, such as SuRF and Rosetta, with the same space usage. We also evaluate SNARF in RocksDB as a filter replacement for filtering requests before they access on-disk data structures. For RocksDB, SNARF can improve the execution time of the system up to 10x compared to SuRF and Rosetta for certain read-only workloads.
Deep learning enables numerous applications across diverse areas. Data systems researchers are also increasingly experimenting with deep learning to enhance data systems performance. We present a tutorial on deep learning, highlighting the data systems nature of neural networks as well as research opportunities for advancements through data management techniques. We focus on three critical aspects: (1) classic design tradeoffs in neural networks which we can enrich through a systems and data management perspective, e.g., thinking critically about storage, data movement, and computation; (2) classic design problems in data systems which we can reconsider with neural networks as a viable design option, e.g., to replace or help system components that make complex decisions such as database optimizers; and (3) essential considerations for responsible application of neural networks in critical human-facing problems in society and how these also link to data management and performance considerations. While these are seemingly a diverse set of rich topics, they are strongly interconnected through data management, and their combination offers rich opportunities for future research.
We present a self-designing key-value storage engine, Cosine, which can always take the shape of the close to "perfect" engine architecture given an input workload, a cloud budget, a target performance, and required cloud SLAs. By identifying and formalizing the first principles of storage engine layouts and core key-value algorithms, Cosine constructs a massive design space comprising of sextillion (1036) possible storage engine designs over a diverse space of hardware and cloud pricing policies for three cloud providers AWS, GCP, and Azure. Cosine spans across diverse designs such as Log-Structured Merge-trees, B-trees, Log-Structured Hash-tables, in-memory accelerators for filters and indexes as well as trillions of hybrid designs that do not appear in the literature or industry but emerge as valid combinations of the above. Cosine includes a unified distribution-aware I/O model and a learned concurrency-aware CPU model that with high accuracy can calculate the performance and cloud cost of any possible design on any workload and virtual machines. Cosine can then search through that space in a matter of seconds to find the best design and materializes the actual code of the resulting storage engine design using a templated Rust implementation. We demonstrate that on average Cosine outperforms state-of-the-art storage engines such as write-optimized RocksDB, read-optimized WiredTiger, and very write-optimized FASTER by 53x, 25x, and 20x, respectively, for diverse workloads, data sizes, and cloud budgets across all YCSB core workloads and many variants.
Everything we see around us is three-dimensional (3D) in nature. The traditional cameras throw light on two-dimensional (2D) behavior of the objects without considering the third dimension or the depth information. This poses a restriction on understanding the complete behavior of any object in space. Thus, this work aims at enlarging the scope of studying and analyzing the 3D behavior of an object through depth imaging. It has its application across the field of robotics and Ergonomics [2]. Ergonomics deals with designing the workspaces and components based on the requirement of the users. The 3D volume occupied by an object moving in space is one of the key focus of ergonomics. This volume representation of an object is known as Motion Envelope (ME). This paper proposes a unique approach to extract the ME of an object moving in space. Techniques like feature extraction is used to detect the required object followed by homography [3] to estimate its pose with motion in space. The 3D position of the object is tracked to reproduce the path followed by the object and a 3D object model is aligned to the detected path to reconstruct the ME. Kinect [6] V1 with colour and depth imaging capabilities is used as the depth sensing device. The results shown in this paper include the object detection and tracking for objects of different shapes and surfaces. The object tracking results are processed to construct ME using the proposed algorithm. This approach emerges to be cheaper and accurate due to the low cost associated with the Kinect sensor compared to the traditional laser-based depth sensors. Thus, we can conclude that this work leads us one-step further in analyzing and understanding the complexity of real-world 3D objects.