Speech-based brain-computer interfaces provide effective voice communication strategies for controlling devices through spoken commands interpreted from brain signals. One of the major challenges in the brain-computer interface problem is the classification of brain signals based on electroencephalography. Electroencephalography is a non-invasive brain signal that is recorded from the scalp surface through electrodes. The signals obtained using easy and cheap equipment have relatively low spatial resolution and high temporal resolution, which requires the most appropriate feature extraction and classification method to achieve optimal results. Also, Collecting sufficient data for a new subject requires a lot of time and effort, so in this paper, data augmentation using a generative adversarial model is proposed to improve the performance of brain signal classification. Also, a model for classifying speech imagery based on brain signals of a new subject with transfer learning using a generative adversarial method based on adversarial domain adaptation method is presented. In order to identify speech imagery, the KaraOne database was used. The proposed method was evaluated with other new methods based on accuracy and kappa criteria. According to the results, the proposed method classifies word imagery and phonemes with 86% and 60.21% accuracy, respectively. The proposed model is independent of each individual's brain signals, which can be effectively classified in the target domain by training the model on the augmented brain signals of the subjects in the source domain without the need for labeled data from the new subject.
In today’s digital world, the vast volume of data generated, often referred to as big data, presents both challenges and opportunities. One significant challenge is the risk of fraud in electronic cash transactions. This study examines and compares 20 common online fraud detection methods within the context of big data, evaluating them based on 11 criteria: type of learning, speed, accuracy, cost (time), complexity, interpretability, scalability, robustness, flexibility, and temporal and spatial complexity. The evaluation highlights the performance of each method against various types of online cash fraud, including identity theft, card skimming, phishing, malware, money laundering, account takeover, refund fraud, and friendly fraud. Performance scores, derived from real-world data and simulations, indicate the effectiveness of each method in identifying and countering fraud in a big data environment. Our findings show that deep learning methods and artificial neural networks outperform other methods in most fraud scenarios, while general rule-based and inferential methods are less effective. This research provides valuable insights for financial institutions, e-commerce platforms, and other online services to enhance their fraud detection capabilities and protect sensitive customer data in the era of big data.
Extracting valuable information from vast sources of social networks while protecting confidentiality and preventing data disclosure is a significant challenge in big data environments. Traditional anonymization methods often fall short in handling the volume, variety, and velocity of big data, leading to high data loss and inefficiency. This article addresses these challenges by proposing a novel anonymization method based on K-means clustering within the Spark framework, leveraging its in-memory processing capabilities. Our model uses K-means clustering to determine optimal cluster heads, significantly reducing data loss and identity disclosure risks. By utilizing Spark's RDD abilities and the MLlib component, our method achieves faster processing times compared to traditional methods that rely on non-in-memory big data tools. Performance evaluation demonstrates that at k = 9, the cost factor is minimized to 0.20, indicating the efficiency and effectiveness of our approach. The proposed method not only enhances processing speed but also ensures minimal data loss, making it suitable for real-time anonymization of big data streams. This work provides a balanced solution that addresses the critical need for high-speed data anonymization while maintaining data privacy and utility.
The rapid advancement of Machine Learning (ML) and Deep Learning (DL) technologies has revolutionized healthcare, particularly in the domains of disease prediction and diagnosis. This study provides a comprehensive review of ML and DL applications across sixteen diverse diseases, synthesizing findings from research conducted between 2015 and 2024. We explore these technologies’ methodologies, effectiveness, and clinical outcomes, highlighting their transformative potential in healthcare settings. Although ML and DL demonstrate remarkable accuracy and efficiency in disease prediction and diagnosis, challenges including quality of data, interpretability of models, and their integration into clinical workflows remain significant barriers. By evaluating advanced approaches and their outcomes, this review not only underscores the current capabilities of ML and DL but also identifies key areas for future research. Ultimately, this work aims to serve as a roadmap for advancing healthcare practices, enhancing clinical decision making, and strengthening patient outcomes through the effective and responsible implementation of AI-driven technologies.
Motor imagery EEG-based brain–computer interfaces (BCIs) have recently been developed for communicating between the brain and external devices. One of the most difficult issues in BCI is classifying the brain activity of motor imagery. To achieve optimal results, the most suitable feature extraction and classification approach must be used because of the poor signal-to-noise ratio, low spatial resolution, and variation in brain activity among subjects in EEG. Gathering enough EEG data takes a lot of time and labor since EEG signals are nonstationary and variable. The main contribution of the paper is to propose models based on deep learning to augment EEG trials to improve EEG classification performance for multi-class motor imagery BCI even when small amounts of EEG data are provided for training. To augment the EEG data in this investigation, generative adversarial models and autoencoders were employed. Then, using various autoencoders, discriminative features were extracted from all the training data and augmented signals. The final step involved classifying multi-class motor imagery of the EEG BCI using deep learning and probabilistic graphical models. For the purpose of classifying EEG signals, the effectiveness of our various feature extraction and data augmentation models was examined using a standard BCI benchmark dataset in terms of classification accuracy, Cohen’s kappa, F1-score, and recall. The suggested model outperforms other cutting-edge models with an average classification accuracy of 90.20
Background: According to the World Health Organization (WHO), approximately 5% of children and 2.5% of adults suffer from attention deficit hyperactivity disorder (ADHD). This disorder can have significant negative consequences on people’s lives, particularly children. In recent years, methods based on artificial intelligence and neuroimaging techniques, such as MRI, have made significant progress, paving the way for development of more reliable diagnostic tools. In this proof of concept study, our aim was to investigate the potential utility of neuroimaging data and clinical information in combination with a deep learning-based analytical approach, more precisely, a novel feature extraction technique for the diagnosis of ADHD with high accuracy. Methods: Leveraging the ADHD200 dataset, which encompasses demographic information and anatomical MRI scans collected from a diverse ADHD population, our study focused on developing modern deep learning-based diagnostic models. The data preprocessing employed a pre-trained Visual Geometry Group16 (VGG16) network to extract two-dimensional (2D) feature maps from three-dimensional (3D) anatomical MRI data to reduce computational complexity and enhance diagnostic power. The inclusion of personal attributes, such as age, gender, intelligence quotient, and handedness, strengthens the diagnostic models. Four deep-learning architectures—convolutional neural network 2D (CNN2D), CNN1D, long short-term memory (LSTM), and gated recurrent units (GRU)—were employed for analysis of the MRI data, with and without the inclusion of clinical characteristics. Results: A 10-fold cross-validation test revealed that the LSTM model, which incorporated both MRI data and personal attributes, had the best diagnostic performance among all tested models in the diagnosis of ADHD with an accuracy of 0.86 and area under the receiver operating characteristic (ROC) curve (AUC) score of 0.90. Conclusions: Our findings demonstrate that the proposed approach of extracting 2D features from 3D MRI images and integrating these features with clinical characteristics may be useful in the diagnosis of ADHD with high accuracy.
Cognitive decision-making processes are crucial aspects of human behavior, influencing various personal and professional domains. This research delves into the application of differential equations in analyzing decision-making accuracy by leveraging eye-tracking data within a virtual industrial town setting. The study unveils a systematic approach to transforming raw data into a differential equation, essential for deciphering the relationship between eye movements during decision-making processes. Mathematical relationship extraction and variable-parameter definition pave the way for deriving a differential equation that encapsulates the growth of fixations on characters. The key factors in this equation encompass the fixation rate (λ) and separation rate (μ), reflecting user interaction dynamics and their impact on decision-making complexities tied to user engagement with virtual characters. For a comprehensive grasp of decision dynamics, solving this differential equation requires initial fixation counts, fixation rate, and separation rate. The formulation of differential equations incorporates various considerations such as engagement duration, character-player distance, relative speed, and character attributes, enabling the representation of fixation changes, speed dynamics, distance variations, and the effects of character attributes. This comprehensive analysis not only enhances our comprehension of decision-making processes but also provides a foundational framework for predictive modeling and data-driven insights for future research and applications in cognitive science and virtual reality environments.
With the growing use of electronic cash cards, the number of transactions with these cards has also increased rapidly, so the importance of using fraud detection models has been paid attention to by financial organizations from various aspects. Electronic cash card fraud detection models often on a single algorithm, optimization of classifications and clusters, to find fraudulent patterns, which provide unsupervised or supervised methods. But the proposed model will use both unsupervised and supervised methods to detect fraud so that it will take advantage of the advantages of both methods. In the proposed method, by selecting the most important features of users’ behavioral patterns such as transaction time and values, their behavioral modeling is done which includes extracting different profiles of users and determining threshold values for each profile. The proposed model will work in real-time by combining two filters to detect electronic card fraud. The first filter is a fast filter that includes a number of unsupervised algorithms, but the second filter is an explicit filter that consists of a number of supervised algorithms. The proposed model creates a profile of the cardholder and measures the degree of deviation of the cardholder’s behavior pattern in new transactions through the Map/Reduce approach for parallel execution alongside the human observer. After then the transaction has been completed and the maximum difference between two consecutive orders of observations, the fraud or non-fraud label will be applied to the transaction and added to the relevant database for future use, in order to detect the deviation of the transaction. According to the simulation results of the proposed model, the accuracy criterion with All Variables, reducing the dimensions of PCA and LASSO is 0.985, 0.987, and 0.980, respectively. F1-Score criterion with All Variables, PCA, and LASSO dimension reduction will be 0.681, 0.676, and 0.669 respectively. The simulation results show that the case with the highest F1 score is the classification of the Proposed Model using all variables. By comparing the simulation results, it can be seen that the F1 score has a high discriminating power between different classification algorithms because the values obtained from it are more different. Also, the results of the calculated performance change values of each pre-processed data with dimension reduction showed that PCA and All variables have very similar performance.
Prostate cancer (PCa) is a prevalent disease among men, with one in eight men experiencing PCa in their lifetime. Predicting the risk of developing metastasis in PCa patients is crucial for clinicians in determining appropriate treatment. This study aimed to identify significant genetic biomarkers for predicting PCa metastasis. The study’s method involved four steps: reading, preprocessing, feature selection, and classification. ANOVA and the ReliefF algorithm were used to select significant features (mRNAs). Additionally, a novel Graph Attention Network (GAT) model based on the Fuzzy theory was introduced, incorporating the Adaptive Neuro-Fuzzy Inference System (ANFIS) for the attention mechanism. The proposed Graph Fuzzy Attention Network (GFAT) model, as part of the proposed classifier, was utilized to classify two groups: PCa patients with and without metastasis. The results showed that the proposed approach achieved an accuracy of 74.4
Big data privacy preservation is a critical challenge for data mining and data analysis. Existing methods for anonymizing big data streams using k-anonymity algorithms may cause high data loss, low data quality, and identity disclosure. In this paper, we propose a novel model for anonymizing big data streams using in-memory processing. The model uses a Spark framework to parallelize the anonymization process and a one-time clustering algorithm to avoid multiple iterations and allocate the data to optimal clusters. We evaluate the performance and effectiveness of the model using a real-world dataset and compare it with three popular k-anonymity algorithms: CRUE, Mean-Shift, and DBSCAN. The results show that the model has the lowest data loss and the highest data quality for different data sizes and k-values. The model is scalable, robust, adaptable, and flexible. The model can provide better data for data mining and data analysis while protecting data privacy and preventing data disclosure.
In the era of big data, with the increase in volume and complexity of data, the main challenge is how to use big data while preserving the privacy of users. This study was conducted with the aim of finding a solution to this challenge. In this study, we examined various data anonymization methods, including differential privacy, advanced encryption, and strong access controls. In addition, the operation, advantages, disadvantages, and use of these methods, the challenges of adapting these methods to big data, and possible solutions for them were also examined. Our results show that traditional data anonymization methods lack scalability, leading to privacy breaches and data loss. When faced with large volumes of data, these methods may not be able to fully process the data. Also, these methods may be ineffective against re-identification attacks, linkage attacks, and inference attacks. We introduced emerging methods that are capable of providing improved privacy with minimal data loss. These methods have scalability for big data. Finally, we examined future research works and raised important questions that can help improve existing algorithms or develop new methods, better manage the complexity and scale of unstructured data.
In light of the escalating privacy risks in the big data era, this paper introduces an innovative model for the anonymization of big data streams, leveraging in-memory processing within the Spark framework. The approach is founded on the principle of K-anonymity and propels the field forward by critically evaluating various anonymization methods and algorithms, benchmarking their performance with respect to time and space complexities. A distinctive formula for optimized cluster determination in the K-means algorithm is presented, along with a novel tuple expiration time strategy for the efficient purging of clusters. The integration of these components into Spark’s RDD and MLlib modules results in a significant decrease in execution time and data loss rates, even with increasing data volumes. The paper’s notable contributions are its methodological advancements that offer a robust, scalable solution for data anonymization, safeguarding user privacy without sacrificing data utility or processing efficiency.
Measuring and predicting accurate joint angles are important to developing analytical tools to gauge users' progress. Such measurement is usually performed in laboratory settings, which is difficult and expensive. So, the aim of this study was continuous estimation of lower limb joint angles during walking using an accelerometer and random forest (RF). Thus, 73 subjects (26 women and 47 men) voluntarily participated in this study. The subjects walked at the slow, moderate, and fast speeds on a walkway, which was covered with 10 Vicon camera. Acceleration was used as input for a RF to estimate ankle, knee, and hip angles (in transverse, frontal, and sagittal planes). Pearson correlation coefficient (r) and Mean Square Error (MSE) were computed between the experimental and estimated data. Paired statistical parametric mapping (SPM) t-test was used to compare the experimental and estimated data throughout gait cycle. The results of this study showed that the MSE of joint angles between the experimental and estimated data ranged from 0.04 to 24.29 and r > 0.91. Moreover, the findings of SPM indicated that there was no significant difference between the experimental and estimated data of ankle, knee, and hip angles in all three planes throughout gait cycle. The results of our research developed a more accessible, portable procedure to quantifying lower limb joint angles by an accelerometer and RF. So, such wearable-based joint angles have the potential to be used in outside-laboratory settings to measure walking kinematics.
Measuring the gait variables outside the laboratory is so important because they can be used to analyze walking in the long run and during real life situations. Wearable sensors like accelerometer show high potential in these applications. So, the aim of this study was continuous estimation of kinetic variables while walking using an accelerometer and artificial neural networks (ANNs). Seventy-three subjects (26 women and 47 men) voluntarily participated in this study. The subjects walked at the slow, moderate, and fast speeds on a walkway which covered with 10 Vicon camera. Acceleration was used as input for a feedforward neural networks to predict the lower limb moments (in sagittal, frontal, and transverse planes), power, and ground reaction force (GRF) (in medial-lateral, anterior-posterior, and vertical directions) during walking. Normalized root mean square error (nRMSE), and Pearson correlation coefficient ( r ) were computed between the measured and predicted variables. Statistical parametric mapping (SPM) was used to compare the measured and predicted variables. The results of this study showed approximately r values of 91–99 and nRMSE values of 4%–15% for GRF, power, and moment between the measured and predicted data. The SPM showed no significant difference between the measured and predicted variables in throughout stance phase. This work has shown the potential of predicting kinetic variables (GRF, moment, and power) in various speeds of walking using the accelerometer. The proposed estimation procedure utilizing a mixture of biomechanics and ANNs can be utilized to solve the tradeoff between richness of data and ease of measuring inherent in wearable sensors.
Machine translation is a technology that reduces costs and speeds up the translation for users by mechanizing translation from one language to another. Machine translation is essentially a step towards globalization for every culture, science, industry, and system. This technology has made great strides in cognitive understanding of natural language since 2013 with the advent of deep learning models. But deep models also always need a lot of data for training. Therefore, the lack of a lot of data and parallel corpus in machine translation is one of the most important problems in this area. Machine translation from English to Persian always suffers from the problem of a lack of resources and data. This article tries to study the deep learning models in machine translation from English to Persian and their strengths and weaknesses, In this article, to solve the problem of lack of English- to-Persian data, the Transformer-based model has been integrated and improved with the Persian language model that has GPT architecture, in addition, the CNN model has been integrated and improved with Autoencoder to improve feature selection and reduce dimensions.
Activation of specific brain areas and synchrony between them has a major role in process of emotions. Nevertheless, impact of anti-synchrony (negative links) in this process still requires to be understood. In this study, we hypothesized that quantity and topology of negative links could influence a network stability by changing of quality of its triadic associations. Therefore, a group of healthy participants were exposed to pleasant and unpleasant images while their brain responses were recorded. Subsequently, functional connectivity networks were estimated and quantity of negative links, balanced and imbalanced triads, tendency to make negative hubs, and balance energy levels of two conditions were compared. The findings indicated that perception of pleasant stimuli was associated with higher amount of negative links with a lower tendency to make a hub in theta band; while the opposite scenario was observed in beta band. It was accompanied with smaller number of imbalanced triads and more stable network in theta band, and smaller number of balanced triads and less stable network in beta band. The findings highlighted that inter regional communications require less changes to receive new information from unpleasant stimuli, although by decrement in beta band stability prepares the network for the upcoming events.
Previous research on fraud detection modeling is often based on a single algorithm, optimizing categories and clusters to find fraudulent patterns that they have provided unsupervised or supervised methods alone and within the framework of Hadoop. The proposed model, a model based on big data analysis extracts important features of user behavior patterns such as time, device type, values, and type of transaction, and their behavioral modeling. By creating different profiles for users, threshold values will be set for each of them. The proposed model for real-time fraud detection of electronic cash payment cards includes two fast and explicit filters. Fast filtering is the combination of the Hidden Markov Model of the first order and Self Organizing Maps (SOM) for fast transaction processing and the Baum-Welch algorithm to find the local Maximum Likelihood. Explicit filter model training is a combination of multilayer Perceptron Neural Network algorithms and Logistic Regression to create a cardholder profile and measure the amount of new transaction deviation from the created profile with the reduction mapping approach or parallel execution of the model. Also, transactions performed for aggregation filters and prediction of the output results of the Schaffer Demister function with the parameter θ in order to detect the degree of deviation of the transaction behavior from the normal state according to the profile of the customer and the maximum difference between two consecutive sequences of observations is used. Finally, the fraudulent or non-fraudulent label on the transaction will be applied and added to the relevant database for storage for later use. According to the model simulation results, the proposed Accuracy, Precision, Recall, and F1-Score criteria are 0.999, 0.9834, 0.7906, and 0.9214, respectively. Simulation results show that the proposed model will perform better in each of the compared criteria against each of the single algorithms.
Agents often need a long time to explore state-action space in order to learn how to act expectedly in Partially Observable Markov Decision Processes (POMDPs). With the reward shaping method, real-time POMDP planning can be guided both in terms of reliability and speed. In this paper, we propose Low Dimensional Policy Graph (LDPG), a new reward shaping method for reducing the dimension of the value function to extract the best state-action pairs. The reward function is then shaped using these key pairs. For accelerating learning speed, we analyze the Transition Function graph to discover significant paths to the learning agent’s goal. Direct comparison on five standard testbeds indicates LDPG brings about the deterministic finding of optimal actions faster regardless of the task type. Our method is shown to reach the goals more quickly (by 41.48 % improvement) and performed 61.57 % better in receiving rewards in the $$ 4 \times 5 \times 2 $$ domain.
The development of diagnostic factors for such neurodegenerative disorders as Alzheimer’s disease is hindered by several limitations that are often witnessed in the early detection of biomarkers in cerebrospinal fluid (CSF) and bio-fluids in clinical contexts. Therefore, here, we suggest a highly sensitive plasmonics nanobiosensor for detection of Alzheimer on one nanodevice through the excitation of surface plasmon resonance (SPR) depending on graphene chemical potential, called a nanobiosensor. This plasmonic nanobiosensor consists of only graphene metasurfaces and samples but does not need extra approaches for the exact analysis of the sample refractive index. Under FDTD simulation, we analyze the sensitivity of 3900 l/RIU for amyloid-beta (Aβ) and figure of merit of 138. This is the first high-performance plasmonics nanobiosensor to monitoring of Alzheimer samples which can be used to detect Alzheimer easily in the future.
Kambiz Badie合作论文数Iran Telecommunication Research Center, Tehran, Iran3