This article presents a Machine Learning Controller (MLC) supported by a P4 switch for improving rate control in non-dedicated Science Demilitarized Zone (Science DMZ) cyberinfrastructures. The proposed scheme utilizes passive data plane measurements such as Round Trip Time (RTT), throughput, queuing delay, and active flow count to regulate campus network output and achieve a desired Data Transfer Node (DTN) target rate. We evaluated our solution through a testbed using a bare-metal data plane switch, legacy router, and emulated hosts. Results show that including a rate controller based on data plane programmable devices on a non-dedicated Science DMZ cyberinfrastructure can effectively improve the completion time of scientific big data flows, while having a low impact on the campus network traffic and bottleneck link utilization. Specifically, the proposed controller achieved an average improvement of 21.72% in Flow Completion Time (FCT) compared to a trivial fixed-rate solution when the DTN uses BBR2 as a Congestion Control Algorithm (CCA). The results highlight the potential of machine learning techniques in conjunction with data plane measurements for optimizing the performance of non-dedicated networks.
Fishing landings in Chile are inspected to control fisheries that are subject to catch quotas. The control process is not easy since the volumes extracted are large and the numbers of landings and artisan shipowners are high. Moreover, the number of inspectors is limited, and a non-automated method is utilized that normally requires months of training. In this work, we propose, design, and implement an automated fish landing control system. The system consists of a custom gate with a camera array and controlled illumination that performs automatic video acquisition once the fish landing starts. The imagery is sent to the cloud in real time and processed by a custom-designed detection algorithm based on deep convolutional networks. The detection algorithm identifies and classifies different pelagic species in real time, and it has been tuned to identify the specific species found in landings of two fishing industries in the Biobío region in Chile. A web-based industrial software was also developed to display a list of fish detections, record relevant statistical summaries, and create landing reports in a user interface. All the records are stored in the cloud for future analyses and possible Chilean government audits. The system can automatically, remotely, and continuously identify and classify the following species: anchovy, jack mackerel, jumbo squid, mackerel, sardine, and snoek, considerably outperforming the current manual procedure.
Fishing has provided mankind with a protein-rich source of food and labor, allowing for the development of an important industry, which has led to the overexploitation of most targeted fish species. The sustainable management of these natural resources requires effective control of fish landings and, therefore, an accurate calculation of fishing quotas. This work proposes a deep learning-based spatial-spectral method to classify five pelagic species of interest for the Chilean fishing industry, including the targeted Engraulis ringens, Merluccius gayi, and Strangomera bentincki and non-targeted Normanichthtys crockeri and Stromateus stellatus fish species. This proof-of-concept method is composed of two channels of a convolutional neural network (CNN) architecture that processes the Red–Green–Blue (RGB) images and the visible and near-infrared (VIS-NIR) reflectance spectra of each species. The classification results of the CNN model achieved over 94% in all performance metrics, outperforming other state-of-the-art techniques. These results support the potential use of the proposed method to automatically monitor fish landings and, therefore, ensure compliance with the established fishing quotas.
This research shows a prototype for crowd location and counting for earthquakes based on deep learning and the infrastructure of a state-of-the-art 5G standalone network deployed at the Universidad de Concepcion, Chile. The system uses an 8 MP panoramic network camera to capture real-time crowd images, which are sent to a Deep Learning Server (DLS) over the 5G network. The camera provides visible color images, and its sensor technology can provide color images even at night. The DLS uses frames from the video feed and generates Focal Inverse Distance Transform (FIDT) maps, in which the counting and location of people are carried out. In particular, the FIDT maps are generated from the crowd images using a deep-learning model composed of two cascaded autoencoders. The 5G technology allows the system to transfer data from the camera to DLS at high speed, an essential feature for a system that will help authorities make critical decisions during natural disasters. Under this scenario, and considering that the number of rescuers is usually limited, our system enables a better distribution of them among several crowded places by instantly knowing the number of people at any time of the day or night.
This paper presents SAFE, a prototype system for supporting the fish landings control of small-scale fishing boats in Chile. SAFE is a modern solution for fishery inspection that automatically discriminates fish species using machine learning. Here, we present a version of SAFE that classifies five target pelagic fish species in Chile: anchovy, Chilean jack mackerel, hake, mote sculpin, and sardine. The system has two stages; the first detects and segments all fish appearing in an image. These segmented images then feed the second stage, which perform species classification. A database of approximately 266 images from these five fish species was constructed for training, validation, and testing purposes. For the fish detection stage, we exploited transfer learning to train Mask R-CNN architectures, an instance segmentation model. As for the fish species classification stage, we exploited transfer learning to train ResNet50 and VGG16 deep learning architectures. Results show that SAFE achieves between 90% and 96.3% macro-average precision (MP) when classifying the five fish species mentioned above. The best architecture, composed of a Mask R-CNN-based detector and a VGG16-based classifier, achieves an MP of 96.3%, which could process a single fish as quick as 16.67 FPS, and one whole 1920x1080-pixel image as quick as 2 FPS.
Many optimal algorithms, heuristics, metaheuristics, simulation approaches, agent-based models, and machine learning tools attempt to solve the job shop scheduling problem (JSSP). This article proposed a model of artificial intelligence with agents representing intelligent products from the perspective of product-driven systems (PDS) to solve this problem at different scales. The intelligent products make all decisions in a distributed way aiming to minimize the makespan and increase the computational efficiency for the JSSP. The agents embed the intelligence function using a based shifting bottleneck heuristic (SBH) approach. The novelty of the proposed approach lies in the automation of decisions in a highly distributed architecture to increase manufacturing flexibility. The results are compared with an optimal integer programming model (IP), SBH, and two conventional heuristics considering instances commonly used in the literature. Concerning the makespan, the proposed approach obtains a fast solution near optimal in instances with a low number of resources and better results than IP and conventional heuristic in instances with a more significant number of resources, increasing the response capacity with a similar computational time.
This paper proposes crowd estimation technology to help authorities make the right decisions in times of crisis. Specifically, deep learning models have faced these challenges, achieving excellent results. In particular, the trend of using single-column Fully Convolutional Networks (FCNs) has increased in recent years. A typical architecture that meets these characteristics is the autoencoder. However, this model presents an intrinsic difficulty: the search for the optimal dimensionality of the latent space. In order to alleviate such difficulty, we propose a dual architecture consisting of two cascaded autoencoders. The first autoencoder is responsible for carrying out the masked reconstruction of the original images, whereas the second obtains crowd maps from the outputs of the first one. In this way, our architecture improves the location of people and crowds in Focal Inverse Distance Transform (FIDT) maps, resulting in more accurate count estimates than estimates obtained through a single autoencoder architecture.
Natural phenomena having catastrophic consequences for people and infrastructure occur every year. Thus, the swift rescue of people is a crucial issue, and the rapid detection of humans trapped in a building can reduce the number of lost lives. Nowadays the use of robots to explore dangerous and inaccessible areas is increasingly common. In such areas, many sensors and actuators are deployed using diverse means to find and rescue people. This paper presents a robotic platform and a set of sensors for exploring an inaccessible area inside a simulated disaster environment. The platform is implemented using the open-source, low-cost Arduino hardware development board. We propose to use information related to carbon dioxide ( $$CO_{2}$$ ) concentrations as a estimate of human breath activity, which, in turn, is used to infer people’s occupancy. Also, we included a contactless thermometer sensor to locate people based on body temperature in order to improve its people detection sensibility.
Earthquakes, and their cascading threats to economic and social sustainability, are a common problem between China and Chile. In such emergencies, automatic image recognition systems have become critical tools for preventing and reducing civilian casualties. Human crowd detection and estimation are fundamental for automatic recognition under life-threatening natural disasters. However, detecting and estimating crowds in scenes is non-trivial due to occlusion, complex behaviors, posture changes, and camera angles, among other issues. This paper presents the first steps i n developing a n intelligent Earthquake Early Warning System (EEWS) between China and Chile. The EEWS exploits the ability of deep learning architectures to properly model different spatial scales of people and the varying degrees of crowd densities. We propose an autoencoder architecture for crowd detection and estimation because it creates compressed representations for the original crowd input images in its latent space. The proposed architecture considers two cascaded autoencoders. The first performs reconstructive masking of the input images, while the second generates Focal Inverse Distance Transform (FIDT) maps. Thus, the cascaded autoencoders improve the ability of the network to locate people and crowds, thereby generating high-quality crowd maps and more reliable count estimates.
Network reliability has become an important concern to network administrators and service providers, and is prominently considered in network design. Particularly, 0-day vulnerabilities are an increasing threat to software-based networking systems. When shared between node appliances, they can be exploited simultaneously and compromise large portions of the network. Moreover, it has been observed that the number of 0-day vulnerabilities discovered yearly in node appliances tends to increase over time. Thus, we can expect that the reliability to 0-day exploits of a network implemented with these appliances will also worsen over time. In this work, we treat network reliability to 0-day exploits as a service, where the network provider agrees to deliver a reliability-based level of service over time. We propose a network reliability metric based on network connectivity and discovered appliance vulnerabilities. We formulate a strategy to guarantee a reliability value over time, based on heterogeneous networking and periodically running cost-effective partial node migrations. We use numerical evaluations to test our methodology on two software-defined wide-area networks based on known backbone IP topologies. Our significant findings are the following: First, when the network reliability becomes worse than the service guarantee, it can be restored in most cases by combining appliance reallocation and node migration. Second, our evaluations show a direct relationship between a network reliability value and the cost incurred to guarantee it. Third, we noted that, when using our appliance-to-node allocation strategy to guarantee the same reliability on different networks, their post-failure connectivity depends on the underlying network topology.
The China-Chile Information and Communication Technology (ICT) Joint Laboratory is a multilateral education-research-production framework for promoting international technical communications, standardization, and industrialization of ICTs between China and Chile. This paper presents the 5G infrastructure built in Chile as part of the Joint Laboratory for collaborative education, innovation, research, and development (R&D). The 5G infrastructure is a compact core network, a New Radio (NR) access network, and a transmission network, which has been implemented as an indoor 5G standalone (SA) architecture operating in the 3.3 to 3.4 GHz band. The core network complies with the 3GPP Release 15 standard and supports three application scenarios defined by the ITU: enhanced mobile broadband, large connections, and low-latency and high-reliability. We also present our experience building the laboratory, the challenges faced for remote commissioning the core system during the COVID-19 pandemics, and how we have engaged and collaborated with the Chilean Government and telecommunications companies for conducting R&D.
A recurrent problem currently affecting network reliability is the simultaneous exploitation of 0-day vulnerabilities shared between several node implementations across the network.When such 0-day vulnerabilities are exploited, large portions of the network may get compromised as a result.In this work, we propose a network node migration strategy to minimize the impact of 0-day attacks on network reliability.The migration method proposes replacing homogeneous node implementations with diverse alternatives to yield a heterogeneous network.The migration method allocates heterogeneous nodes within the network by minimizing the product between the average and the maximum number of network partitions, which may emerge after the simultaneous exploitation of 0-day risks on shared network resources.As we show, our migration strategy maximizes network connectivity in the event of a simultaneous 0-day attack.Our work's significant findings are the following: First, increasing the heterogeneity in node technologies reduces the attacker's ability to break down the entire network.Second, given a set of available network technologies that partially share risks, a network design implemented using several heterogeneous technologies sharing a small number of 0-day risks is more reliable than one with a small number of technologies whose 0-day risks are disjoint.Third, we observed that in a node-heterogeneous network topology, clustering nodes by technology improves network reliability.
Biomedical text classification algorithms, which currently support clinical decision-making processes, call for expensive training texts due to the low availability of labeled corpus and the cost of manual annotation by specialized professionals. The active learning (AL) approach to classification heavily lessens such cost by reducing the number of labeled documents required to achieve specified performance. This article introduces a query strategy and a stopping criterion that transform CREGEX, a regular-expressions-based text classification algorithm, in an AL biomedical text classifier. The query strategy samples the training dataset, trading off the greedy learning achieved by the regular expressions classification precision and the conservative learning induced by text sequence alignment classification. The sustained reduction in the variance of the query strategy scores is used as a stopping criterion. The AL classifier was compared with Support Vector Machine (SVM), Naïve Bayes (NB), and a classifier based on Bidirectional Encoder Representations from Transformers (BERT), using three datasets with biomedical information in Spanish on smoking habits, obesity, and obesity types. The learning curve results indicate that AL in CREGEX allowed to efficiently reduce the number of training examples for equal performance than the rest of the classifiers, obtaining areas under the learning curve greater than 85% in all cases. The stopping criterion applied to the AL process allowed to use, on average, approximately 32% to 50% of the total training examples with differences in performance concerning the maximum value of the learning curve not exceeding 2%. This performance demonstrates the effectiveness of using AL in a biomedical text classifier based on regular expressions, which is attributable to such expressions' ability to represent intricate sequential patterns in training texts considered most informative.
Although the importance of router’s buffer sizing in network performance is well known, estimating the current size of the bottleneck buffer is an open research problem. This paper presents a method to achieve such estimation, for the case where the bottleneck buffer operates under a finite number of buffer sizing regimes. The scheme uses a supervised machine learning approach to properly model such regimes and a classification mechanism to predict the coarse buffer size using the following end-to-end network measurements, which are collected at the sender side: throughput, Round Trip Time (RTT), and Congestion Window (CWND). In contrast to previous work, the scheme does not assume a homogeneous congestion control algorithm used by the senders. The proposed approach was tested using data collected on a real testbed. The corresponding results show that the Support Vector Machine (SVM) Radial Basis Function (RBF) classifier correctly estimates the bottleneck buffer size, under different network conditions.
High accuracy text classifiers are used nowadays in organizing large amounts of biomedical information and supporting clinical decision-making processes. In medical informatics, regular expression-based classifiers have emerged as an alternative to traditional, discriminative classification algorithms due to their ability to model sequential patterns. This article presents CREGEX (Classifier Regular Expression), a biomedical text classifier based on an automatically generated regular-expressions-based feature space. We conceived an algorithm for automatically constructing an informative and discriminative regular-expressions-based feature space, suitable for binary and multiclass discrimination problems. Regular expressions are automatically generated from training texts using a coarse-to-fine text aligning method, which trades off the lexical variants of words, in terms of gender and grammatical number, and the generation of a feature space containing a large number of noisy features. CREGEX carries out feature selection by filtering keywords and also computes a confidence metric to classify test texts. Three de-identified datasets in Spanish, with information on smoking habits, obesity, and obesity types, were used here to assess the performance of CREGEX. For comparison, Support Vector Machine (SVM) and Naïve Bayes (NB) supervised classifiers were also trained with consecutive sequences of tokens (n-grams) as features. Results show that, in all the datasets used for evaluation, CREGEX not only outperformed both the SVM and NB classifiers in terms of accuracy and F-measure (p-value<; 0.05) but also used a fewer amount of training examples to achieve the same performance. Such a superior performance is attributed to the regular expressions' ability to represent complex text patterns.
In this paper, we propose a heterogeneous risk-aware Software-Defined Networking (SDN) migration method for designing survivable networks in the face of multiple correlated failures. The migration method, which is implemented in one shot, specifies how many nodes of each SDN implementation are needed, and where such nodes must be located, in order to yield an SDN migrated network with maximal survivability, when multiple correlated failures impact the entire network connectivity. We formulated the survivable SDN migration problem through integer optimization, where the proposed cost function assesses the survivability of the migrated network in terms of the number of connected components after a failure. The numerical results calculated over test networks show the capability of migration method to provide survivable SDN topologies, which trade-off the heterogeneity in the SDN implementations and the number of shared risks.
In this work, we present FREGEX a method for automatically extracting features from biomedical texts based on regular expressions. Using Smith-Waterman and Needleman-Wunsch sequence alignment algorithms, tokens were extracted from biomedical texts and represented by common patterns. Three manually annotated datasets with information on obesity, obesity types, and smoking habits were used to evaluate the effectiveness of the proposed method. Features extracted using consecutive sequences of tokens (n-grams) were used for comparison, and both types of features were mathematically represented using the TF-IDF vector model. Support Vector Machine and Naïve Bayes classifiers were trained, and their performances were ultimately used to assess the ability of the feature extraction methods. Results indicate that features based on regular expressions not only improved the performance of both classifiers in all datasets but also use fewer features than n-grams, especially in those datasets containing information related to anthropometric measures (obesity and obesity types).
Current data networks are highly homogeneous because of management, economic, and interoperability reasons. This technological homogeneity introduces shared risks, where correlated failures may entirely disrupt the network operation and impair multiple nodes. In this paper, we tackle the problem of improving the resilience of homogeneous networks, which are affected by correlated node failures, through optimal multiculture network design. Correlated failures regarded here are modeled by SRNG events. We propose three sequential optimization problems for maximizing the network resilience by selecting as different node technologies, which do not share risks, and placing such nodes in a given topology. Results show that in the 75% of real-world network topologies analyzed here, our optimal multiculture design yields networks whose probability that a pair of nodes, chosen at random, are connected is 1, i.e., its ATTR metric is 1. To do so, our method efficiently trades off the network heterogeneity, the number of nodes per technology, and their clustered location in the network. In the remaining 25% of the topologies, whose average node degree was less than 2, such probability was at least 0.7867. This means that both multiculture design and topology connectivity are necessary to achieve network resilience.
Natural disasters, depending on both how many occur concurrently and their size, may produce large-scale correlated failures in data network infrastructure. These failures may cause service interruptions due to disconnections of nodes in the network. Proper fault modeling is crucial to calculate network damage, determine which data paths will remain active between a pair of nodes, and thus maintain a resilient network. While in the literature different sizes of circular shapes are used to model fault regions, in this work a new fault model is proposed. The model adjusts to the granularity level established by the network opera- tor to define the size and number of concurrent fault regions. Equipped with the failure model, it is possible to observe, through disjoint paths problem, the advantages of using micro failure region models to mitigate false positive failures associated when macro failure region is used.
Software-defined networking (SDN) technology is being widely adopted across commercial, governmental, and educational networking domains. However, teaching SDN-related concepts is a challenge due to the inherent nature of this paradigm, its detailed technological tools/skill sets, and the continual evolution of the information and communications technology (ICT) sector. As a result, there is a growing need to support new initiatives to develop applied research and hands-on training methodologies. Accordingly, this article presents a framework that has been developed/evolved over the past few years for teaching and conducting research activities related to SDN. The key idea here is to leverage the close relationship between SDN, real-world ICT problems, and innovative services development. Thus, the proposed framework integrates actors, stakeholders, phases, and components, along with their interrelationships in an applied research environment. Some related research and development experiences from the application of this framework in Chile are also presented to highlight its contributions toward creating a wider knowledge base and benefiting society.