Public authorities and private companies have used video cameras as part of surveillance systems, and one of their objectives is the rapid detection of physically violent actions. This task is usually performed by human visual inspection, which is labor-intensive. For this reason, different deep learning models have been implemented to remove the human eye from this task, yielding positive results. One of the main problems in detecting physical violence in videos is the variety of scenarios that can exist, which leads to different models being trained on datasets, leading them to detect physical violence in only one or a few types of videos. In this work, we present an approach for physical violence detection on images obtained from video based on threshold active learning, that increases the classifier’s robustness in environments where it was not trained. The proposed approach consists of two stages: In the first stage, pre-trained neural network models are trained on initial datasets, and we use a threshold (μ) to identify those images that the classifier considers ambiguous or hard to classify. Then, they are included in the training dataset, and the model is retrained to improve its classification performance. In the second stage, we test the model with video images from other environments, and we again employ (μ) to detect ambiguous images that a human expert analyzes to determine the real class or delete the ambiguity on them. After that, the ambiguous images are added to the original training set and the classifier is retrained; this process is repeated while ambiguous images exist. The model is a hybrid neural network that uses transfer learning and a threshold μ to detect physical violence on images obtained from video files successfully. In this active learning process, the classifier can detect physical violence in different environments, where the main contribution is the method used to obtain a threshold μ (which is based on the neural network output) that allows human experts to contribute to the classification process to obtain more robust neural networks and high-quality datasets. The experimental results show the proposed approach’s effectiveness in detecting physical violence, where it is trained using an initial dataset, and new images are added to improve its robustness in diverse environments.
The global emergency of COVID-19 caused by the SARS-CoV-2 virus at the end of 2019, was without a doubt a critical and historical point for society in general; for instance, the effective development of vaccines, as well as the efficient distribution of them; They were an unprecedented challenge to slow down the spread or mitigate its impact on societies around the world. This article specifies three bio-inspired metaheuristic algorithms (genetic algorithm, particle swarm optimization algorithm, and artificial bee colony algorithm) that were used in the context of the capacitated vehicle routing problem to generate vaccine distribution routes, specifically, COVID-19 vaccine for over 18 years old the first and the second doses applications in Mexico, particularly in the State of Mexico. The quality of the solutions obtained by these algorithms is compared, as a result of the performance of the particle swarm optimization (PSO) algorithm being superior in solution quality.
Violence against women captured in videos and surveillance systems necessitates effective identification to enable appropriate reactions for controlling and mitigating of its effects in public spaces and the potential apprehension of aggressors. While several algorithms have been developed for violence detection, their evaluation has primarily focused on controlled scenarios with clear differentiation between violent and non-violent scenes, representing two-class identification problems. However, real-world situations often present challenges where specific actions, such as hugs or effusive greetings, fall into an ambiguous class that is difficult to classify. Consequently, this transforms into a multi-class identification problem. In this study, we assess the performance of three pre-trained models, namely VGG16, ResNet50, and InceptionV3, to evaluate their efficacy in addressing the multi-class identification challenges. Furthermore, we compare their performance against datasets consisting of two-class classifications, where the models generally exhibit satisfactory results. Our analysis reveals that the models struggle to differentiate the ambiguous scenes effectively, with Inception V3 achieving a 0
An innovative strategy for organizations to obtain value from their large datasets, allowing them to guide future strategic actions and improve their initiatives, is the use of machine learning algorithms. This has led to a growing and rapid application of various machine learning algorithms with a predominant focus on building and improving the performance of these models. However, this data-centric approach ignores the fact that data quality is crucial for building robust and accurate models. Several dataset issues, such as class imbalance, high dimensionality, and class overlapping, affect data quality, introducing bias to machine learning models. Therefore, adopting a data-centric approach is essential to constructing better datasets and producing effective models. Besides data issues, Big Data imposes new challenges, such as the scalability of algorithms. This paper proposes a scalable hybrid approach to jointly addressing class imbalance, high dimensionality, and class overlapping in Big Data domains. The proposal is based on well-known data-level solutions whose main operation is calculating the nearest neighbor using the Euclidean distance as a similarity metric. However, these strategies may lose their effectiveness on datasets with high dimensionality. Hence, the data quality is achieved by combining a data transformation approach using fractional norms and SMOTE to obtain a balanced and reduced dataset. Experiments carried out on nine two-class imbalanced and high-dimensional large datasets showed that our scalable methodology implemented in Spark outperforms the traditional approach.
In Latin American and Caribbean States the verbal violence against women on social networks, such as Twitter, is a serious threat that has been addressed through the implementation of social norms, public policies, and social movements. Nevertheless, a challenge is the effective and automatic real-time detection of violent tweets. In this sense, traditional machine learning algorithms have been proposed to tackle social issues where the training process is performed in a static manner. However, considering that Twitter is a dynamic environment where a vast of tweets are generated each second, it requires powerful machine learning algorithms that could exploit this pool of unlabeled data to be incorporated into the model through continuous updates. This paper explores an active learning method based on uncertainty sampling, which identifies the most confusing tweets to be labeled by an expert in real-time. This focused selection prioritizes which data can be used to train a multilayer perceptron that can achieve a better performance with fewer training samples. Experimental results show that including new samples yields promising results, increasing the AUC from 0.8712 to 0.8833.
The interest in exploiting big datasets with machine learning has led to adapting classic strategies in this new paradigm determined by volume, speed, and variety. Because data quality is a determining factor in constructing a classifier, it has also been necessary to adapt or develop new data preprocessing techniques. One of the challenges of most significant interest is the class imbalance problem, where the class of interest has a smaller number of examples concerning another class called the majority. To alleviate this problem, one of the most recognized techniques is SMOTE, which is characterized by generating instances of the minority class through a process that uses the nearest neighbor rule and the Euclidean distance. Various articles have shown that SMOTE is not appropriate for datasets with high dimensionality. However, in big data, datasets with high dimensionality have contained many zeros. Therefore, in this article, our objective is to analyze the SMOTE-BD behavior on imbalanced big datasets with sparse and dense dimensionality. Experimental results using two classifiers and big datasets with different dimensionalities suggest that sparsity is a predominant factor than the dimensionality in the behavior of SMOTE-BD.
Artificial Neural Networks (ANN) have encountered interesting applications in forecasting several phenomena, and they have recently been applied in understanding the evolution of the novel coronavirus COVID-19 epidemic. Alone or together with other mathematical, dynamical, and statistical methods, ANN help to predict or model the transmission behavior at a global or regional level, thus providing valuable information for decision-makers. In this research, four typical ANN have been used to analyze the historical evolution of COVID-19 infections in Mexico: Multilayer Perceptron (MLP), Convolutional Neural Networks (CNN), Long Short-Term Memory (LTSM) neural networks, and the hybrid approach LTSM-CNN. From the open-source data of the Resource Center at the John Hopkins University of Medicine, a comparison of the overall qualitative fitting behavior and the analysis of quantitative metrics were performed. Our investigation shows that LSTM-CNN achieves the best qualitative performance; however, the CNN model reports the best quantitative metrics achieving better results in terms of the Mean Squared Error and Mean Absolute Error. The latter indicates that the long-term learning of the hybrid LSTM-CNN method is not necessarily a critical aspect to forecast COVID-19 cases as the relevant information obtained from the features of data by the classical MLP or CNN.
The pandemic caused by the COVID-19 disease has affected all aspects of the life of the people in every region of the world. The academic activities at universities in Mexico have been particularly disturbed by two years of confinement; all activities were migrated to an online modality where improvised actions and prolonged isolation have implied a significant threat to the educational institutions. Amid this pandemic, some opportunities to use Artificial Intelligence tools for understanding the associated phenomena have been raised. In this sense, we use the K-means algorithm, a well-known unsupervised machine learning technique, to analyze the data obtained from questionaries applied to students in a Mexican university to understand their perception of how the confinement and online academic activities have affected their lives and their learning. Results indicate that the K-means algorithm has better results when the number of groups is bigger, leading to a lower error in the model. Also, the analysis helps to make evident that the lack of adequate computing equipment, internet connectivity, and suitable study spaces impact the quality of the education that students receive, causing other problems, including communication troubles with teachers and classmates, unproductive classes, and even accentuate psychological issues such as anxiety and depression.
The problem of gender-based violence in Mexico has been increased considerably. Many social associations and governmental institutions have addressed this problem in different ways. In the context of computer science, some effort has been developed to deal with this problem through the use of machine learning approaches to strengthen the strategic decision making. In this work, a deep learning neural network application to identify gender-based violence on Twitter messages is presented. A total of 1,857,450 messages (generated in Mexico) were downloaded from Twitter: 61,604 of them were manually tagged by human volunteers as negative, positive or neutral messages, to serve as training and test data sets. Results presented in this paper show the effectiveness of deep neural network (about 80% of the area under the receiver operating characteristic) in detection of gender violence on Twitter messages. The main contribution of this investigation is that the data set was minimally pre-processed (as a difference versus most state-of-the-art approaches). Thus, the original messages were converted into a numerical vector in accordance to the frequency of word's appearance and only adverbs, conjunctions and prepositions were deleted (which occur very frequently in text and we think that these words do not contribute to discriminatory messages on Twitter). Finally, this work contributes to dealing with gender violence in Mexico, which is an issue that needs to be faced immediately.
Clustering algorithms have been used in different areas of knowledge with different goals such as noise detection, outliers, and descriptive tasks. The adsorption kinetics is a curve that describes the rate retention to the adsorbate on the adsorbent at time, which is represents as a two-dimensional graph. In this paper, we present a computational application to determine the experimental conditions that influence when equilibrium point is reached into adsorption kinetics curve using the K-means clustering algorithm and, Parallel Coordinates concept, in order to prove our method we used adsorption kinetic curves Q-PVA . Results obtained were compared with two designs of experiments (three-stage nested design and hierarchical design with crossed factors).
During COVID-19 quarantine, in online sites such as social networks, Gender-Based Violence has alarmingly increased. Online platforms have taken various measures to regulate and prevent broadcasting of violence messages. Multiple proposals based on machine learning and deep learning approaches have been used to address this problem. This work presents an improvement in implementation of a deep learning neural network for detection of Gender-Based Violence in Twitter messages. A total of 32,500 tweets were downloaded from Mexican Twitter accounts and human volunteers manually tagged the tweets as violent and non-violent to be used as training and testing data sets. Experimental results show the effectiveness of the deep neural network (about 90% of the Area Under the Receiver Operating Characteristic) to detect gender violence in Twitter messages using a simple Natural Language Processing approach.
This work reports DFT calculations for the assessment of metallic decoration of boron substitution Zeolite Templated Carbon vacancy for hydrogen adsorption. The boron substitution on Zeolite Templated Carbon vacancy is characterized by the formation of pentagonal and heptagonal rings. Moreover, the boron substitution can be considered as a promising way for hydrogen storage, this way boron substitution is used on Zeolite Templated Carbon vacancy in order to create an active site for metallic decoration. Once that we develop a Boron substitution on Zeolite Templated Carbon vacancy, the decoration with Lithium, Sodium, and Calcium atoms is also carried out. The analysis reveals that the Na decoration has the best performance for hydrogen storage. The results show that boron substitution on Zeolite Templated Carbon vacancy decorated with 3 Sodium atoms can adsorb up to fifteen hydrogen molecules (5 hydrogen molecules per Sodium atom), this gives a gravimetric storage capacity of 6.55 % wt., which is enough for meeting DOE gravimetric targets. In addition, the average binding energies and adsorption energies are calculated in the range 0.2298-0.2144 eV/H-2, which constitute desirable energies for hydrogen adsorption. Besides, the hydrogen adsorption process is carried out by electrostatic interaction between the Na cation and the induced H-2 dipole. The calculation performed in this work reveals that the boron substitution on Zeolite Templated Carbon vacancy decorated with Na atoms is a good candidate as a medium for hydrogen storage. (C) 2020 Hydrogen Energy Publications LLC. Published by Elsevier Ltd. All rights reserved.
Let $$G=(V, E)$$ be a graph with a vertex set V and set of edges E. The Graph Coloring Problem consists of splitting the set V into k independent sets (color classes); if two vertices are adjacent (i.e. vertices which share an edge), then they cannot have the same color. In order to address this problem, a plethora of techniques have been proposed in literature. Those techniques are especially based on heuristic algorithms, because the execution time noticeably increases if exact solutions are applied to graphs with more than 100 vertices. In this research, a metaheuristic approach that combines a deterministic algorithm and a heuristic algorithm is proposed, in order to approximate the chromatic number of a graph. This method was experimentally validated by using a collection of graphs from the literature in which the chromatic number is well-known. Obtained results show the feasibility of the metaheuristic proposal in terms of the chromatic number obtained. Moreover, when the proposed methodology is compared against robust techniques, this procedure increases the quality of the residual graph and improves the Tabu search that solves conflicts involved in a path as coloring phase.
Earthquakes are events that cannot be predicted. However, when they occur, devastating consequences are shown in economic, social and structural areas, among others. In this paper, the mining of association rules is carried out in order to estimate the repair cost required by schools affected during the earthquakes of September 7th and 19th, of 2017 in Mexico. For that, we use the public data collected by the Mexican FONDEN.
The class imbalance problem has been a hot topic in the machine learning community in recent years. Nowadays, in the time of big data and deep learning, this problem remains in force. Much work has been performed to deal to the class imbalance problem, the random sampling methods (over and under sampling) being the most widely employed approaches. Moreover, sophisticated sampling methods have been developed, including the Synthetic Minority Over-sampling Technique (SMOTE), and also they have been combined with cleaning techniques such as Editing Nearest Neighbor or Tomek’s Links (SMOTE+ENN and SMOTE+TL, respectively). In the big data context, it is noticeable that the class imbalance problem has been addressed by adaptation of traditional techniques, relatively ignoring intelligent approaches. Thus, the capabilities and possibilities of heuristic sampling methods on deep learning neural networks in big data domain are analyzed in this work, and the cleaning strategies are particularly analyzed. This study is developed on big data, multi-class imbalanced datasets obtained from hyper-spectral remote sensing images. The effectiveness of a hybrid approach on these datasets is analyzed, in which the dataset is cleaned by SMOTE followed by the training of an Artificial Neural Network (ANN) with those data, while the neural network output noise is processed with ENN to eliminate output noise; after that, the ANN is trained again with the resultant dataset. Obtained results suggest that best classification outcome is achieved when the cleaning strategies are applied on an ANN output instead of input feature space only. Consequently, the need to consider the classifier’s nature when the classical class imbalance approaches are adapted in deep learning and big data scenarios is clear.
Fourteen formulas from the state-of-art were used in this paper to find the optimal number of neurons in the hidden layer of an autoencoder neural network. The latter is employed to reduce the dataset dimension on high-dimensionality scenarios with not significant reduction in classification accuracy in comparison to the use of the whole dataset. A Deep Learning neural network was employed to analyze the effectiveness of the studied formulas in classification terms (accuracy). Eight high-dimensional datasets were processed in an experimental set in order to assess this proposal. Results presented in this work show that formula 13 (used to find the number of hidden neurons of the auto-encoder) is effective to reduce the data dimensionality without a statistically significant reduction of the classification performance, as it was confirmed by the Freidman test and the Holm's post-hoc test.
The class imbalance problem is a challenging situation in machine learning but also it appears frequently in recent Big Data applications. The most studied techniques to deal with the class imbalance problem have been Random Over Sampling (ROS), Random Under Sampling (RUS) and Synthetic Minority Over-sampling Technique (SMOTE), especially in two-class scenarios. However, in the Big Data scale, multi-class imbalance scenarios have not extensively studied yet, and only a few investigations have been performed. In this work, the effectiveness of ROS and SMOTE techniques is analyzed in the Big data multi-class imbalance context. The KDD99 dataset, which is a popular multi-class imbalanced big data set, was used to probe these oversampling techniques, prior to the application of a Deep Learning Multi-Layer Perceptron. Results show that ROS and SMOTE are not always enough to improve the classifier performance in the minority classes. However, they slightly increase the overall performance of the classifier in comparison to the unsampled data.
In this work, we report DFT calculations of the energy formation and stability of multivacancies in a unit of Zeolite Template Carbon (C39H9). We label as V-n the respective vacancy where n carbon atoms have been removed from the pristine C39H9 structure. The results show that V-2, V-4, V-6 and V-9 are the most stable vacancies on the ZTC structure. This result agrees with many other studies. Besides, the most stable vacancy of ZTC structure is when nine carbon atoms are removed (V-9) from the ZTC structure. The formation of pentagon rings in the reconstruction of the ZTC vacancy give drastic effect on the energetics stability. Therefore, the formation of pentagon rings eliminates the dangling bonds thus lowering the energy formation. It is also carried out the decoration of ZTC vacancy with Lithium and Calcium atoms, this is the way to use de ZTC vacancy decorated as a medium for hydrogen storage. The results show that the ZTC vacancy decorated with 3 Lithium atoms can adsorb a maximum of nine hydrogen molecules (3 hydrogen molecules per Lithium atom). This gives a gravimetric storage capacity of 4.44 wt percent (wt. %), which is not enough for meeting DOE gravimetric target. On the other hand, to reach DOE gravimetric target, the study of ZTC vacancy decorated with 3 Calcium atoms is carried out, which can adsorb maximum of fifteen hydrogen molecules (5 hydrogen molecules per Calcium atom), this gives gravimetric storage capacity of 5.81 wt %, which meet DOE gravimetric targets, besides the binding energy of hydrogen molecules on ZTC vacancy decorated with 3 Calcium is calculated. These energies are in the range 0.2453-0.2053 eV/H-2, which are desirable energies for hydrogen adsorption. This is demonstrated by building isotherm adsorption path. The results show that forming vacancies on ZTC structure decorated with three Calcium atoms (3Ca-C30H9) is good candidate as medium for hydrogen storage. (C) 2019 Hydrogen Energy Publications LLC. Published by Elsevier Ltd. All rights reserved.
Neural networks methodology is a tool that allows to get the potential energy curve in cases where the data dispersion does not fit a discrete distribution; hence, a binding energy fitting can be found with this methodology. A data distribution of the intermolecular pair interaction potential in vacuum has been previously accomplished between asphaltene-asphaltene (U-AA) systems by using compass classical force field. In the latter, all possible interaction geometries are taken into account between the species: random, face-to-face, t-shape and edge to edge. In one of these cases, a potential energy curve is gotten when the geometry of interaction is face-to-face using a statistical fit. Focusing in these data distribution, neural networks have been applied on the following cases: i) face-to-face distribution of asphaltene-asphaltene interactions; ii) the complete asphaltene-asphaltene discrete distribution of energy vs contact distance (the minimum distance at which the interacting species is not equal to zero) where all-geometries were used, and iii) the random distribution of geometries of asphaltene-asphaltene interactions. In addition, using an asphaltene model molecule reported by Speight and taking into account two possible asphaltene interactions (face-to-face and random), firstly the data distribution of energy as a function of distance is obtained, and secondly neural networks are applied to fit the corresponding potential energy curve.
On-line learning is a training paradigm that allows the processing of constant data flows, so that learning adapts to new knowledge. However, due to the nature of the study problem, it is possible that in the clustering obtained there are data complexities (outliers, atypical patterns, noisy, etc.) that deteriorate the performance of the model in the classification stage. Due to the above, an alternative to cope data complexities is the use of algorithms that allow to detect reject options to filter noisy pattern. In this research the neighborhood-based reject option is implemented in an on-line learning process, with the intention of improving the clustering quality and thus increasing the precision indexes obtained with the nearest neighbor's rule in the classification stage. Likewise, to validate the quality of the clustering generated, internal and external analysis metrics are used. The experimental results show the viability of the proposal when analyzed on real data.