We consider the task of reconstructing the cabling arrangements of last-mile telecommunication networks using customer modem data. In such networks, downstream data traverses from a source node down through the branches of the tree network to a set of customer leaf nodes. Each modem monitors the quality of received data using a series of continuous data metrics. The state of the data, when it reaches a modem, is contingent upon the path it traverses through the network and can be affected by, e.g., corroded cable connectors.We train an encoder to identify irregular inherited events in modem quality data, such as network faults, and encode them as discrete data sequences for each modem. Specifically, the encoding scheme is obtained by using unsupervised contrastive learning, where a Siamese neural network is trained on a positive (true) topology, its modem data, and a set of negative (false) topologies. The weights of the Siamese network are continuously updated based on a new modified version of the Maximum Parsimony optimality criterion. This approach essentially integrates an optimization problem directly into a deep learning loss function.We evaluate the encoder’s performance on simulated data instances with randomly added events. The performance of the encoder is tested both on its ability to extract and encode events as well as whether the encoded data sequences lead to accurate topology reconstructions under the modified version of the Maximum Parsimony optimality criterion.Promising computational results are reported for trees with a varying number of internal nodes, up to a maximum of 20. The encoder identifies a high percentage of simulated events, leading to nearly perfect topology reconstruction. Overall, these results affirm the potential of embedding an optimization problem into a deep learning loss function, unveiling many interesting topics for further research.
Hybrid Fiber-Coaxial (HFC) networks are a popular infrastructure for delivering internet to consumers, however, they are complex and susceptible to various errors. Internet service providers currently rely on manual operations for network monitoring, underscoring the need for automated fault detection. We propose a novel framework for estimating the density of multivariate time series, tailored for anomaly detection in broadband networks. Our framework comprises two phases. In the first phase, we employ an autoencoder based on one-dimensional convolutions to learn a latent representation of time series windows, thereby preserving context. In the second phase, we utilize a Normalizing Flow (NF) to model the distribution within this latent space, enabling subsequent anomaly detection. For efficient separation, we propose an iterative weighing algorithm allowing the NF to model only the systematic behavior, thereby separating outlying behavior. We validated our methodology using a publically available synthetic dataset and real-world data from TDC NET, Denmark’s leading provider of digital infrastructure. Initial experiments with the synthetic dataset demonstrated that our density-based estimator effectively distinguishes anomalies from normal behavior. When applied to the unlabeled TDC NET dataset, our framework exhibits promising performance, identifying outliers clustering themselves away from the high-density region, thus enabling subsequent root cause analysis.
Monitoring of flocculation processes such as those used in downstream processing of a fermentation broth is essential for process control. One approach is to apply microscopic imaging combined with image analysis for characterizing the state of the process. In this work, we investigate and compare the use of supervised feedforward convolutional neural network (CNN) architectures to predict the process states from the image information and compare the results with the traditional alternative of characterizing flocs based on manually engineered image features guided by human expertise. From a well-defined image data set representing six process states, the objective is to establish end-to-end classification models which are accurate but at the same time learn meaningful latent variable space representations. Specifically, we evaluate three different CNN architectures with varying degrees of regularization and compare results with logistic regression models based on inputs from two different traditional feature engineering methods. By applying global average pooling as a structural regularizer to the CNN architecture, we significantly improve the generalization performance in comparison with the classification accuracies of the traditional feature engineered models. Furthermore, we show that by imposing a projection to latent structures (PLS) like regularization framework onto the CNN, it can also learn a latent variable representation that mimics the features selected by human expertise. This work explores the use of convolutional neural network (CNN) to monitor industrial flocculation processes, comparing them with traditional methods guided by human expertise. By evaluating different CNN architectures and traditional feature engineering methods, we aim to establish accurate classification models, while learning meaningful and interpretable latent variable space representations.
Convolutional Neural Networks (CNN) are the state-of-the-art in the field of visual computing. However, a major problem with CNNs is the large number of floating point operations (FLOPs) required to perform convolutions for large inputs. When considering the application of CNNs to video data, convolutional filters become even more complex due to the extra temporal dimension. This leads to problems when respective applications are to be deployed on mobile devices, such as smart phones, tablets, micro-controllers or similar, indicating less computational power. Kim et al. proposed using a Tucker-decomposition to compress the convolutional kernel of a pre-trained network for images in order to reduce the complexity of the network, i.e. the number of FLOPs. In this paper, we generalize the aforementioned method for application to videos (and other 3D signals) and evaluate the proposed method on a modified version of the THETIS data set, which contains videos of individuals performing tennis shots. We show that the compressed network reaches comparable accuracy, while indicating a memory compression by a factor of 51. However, the actual computational speed-up (factor 1.4) does not meet our theoretically derived expectation (factor 6).
Endo-fucoidanases, including EC 3.2.1.211 endo-alpha-1,3-L-fucanase and EC 3.2.1.212 endo-alpha-1,4-L-fucanase activities, catalyze depolymerization of fucoidans - a group of bioactive, sulfated fucosyl-polysaccharides found primarily in brown macroalgae (brown seaweeds). Quantitative assessment of endo-fucoidanase activity is critical for characterizing endo-fucoidanase kinetics and for comparing the action of different endo-fucoidanases on different types of fucoidans. However, the current state-of-the-art endo-fucoidanase assay consists of a qualitative assessment based on Carbohydrate-Polyacrylamide Gel Electrophoresis. Here, we report a new quantitative endo-fucoidanase assay based on real time spectral evolution profiling of changes in substrate and product during endo-fucoidanase action using Fourier Transform InfraRed spectroscopy (FTIR) combined with Parallel Factor Analysis (PARAFAC). The FTIR-PARAFAC assay was validated by monitoring the reaction progress of three different microbial endo-fucoidanase enzymes, FcnA Delta 229, FFA2 and Fhf1 Delta 470, on two different fucoidan substrates. The substrates were purified from the brown macroalgae Fucus evanescens and Fucus vesiculosus, respectively. The evolution profiling showed that the strongest spectral change of the fucoidans during enzymatic depolymerization occurred in the spectral range 1220-1260 cm(-1), but the profiles differed depending on the substrate and the enzyme used. Spectral changes within 1220-1260 cm(-1) are in agreement with the enzymatic depolymerization inducing signature changes in the mid-infrared absorption of sulfated fucosyls as sulfate ester bonds and C-O stretching vibrations absorb in this spectral region. Based on the data obtained, we also introduce an activity unit for endo-fucoidanases: One endo-fucoidanase Unit, Uf, is the amount of enzyme able to catalyze a change in the FTIR-PARAFAC score by 0.01 during 498 s of reaction (8.3 min) on 20 g/L pure fucoidan from F. evanescens at 42 degrees C, pH 7.4, 100 mM NaCl and 10 mM CaCl2. This new quantitative endo-fucoidanase assay can pave the way for better kinetic characterizations as well as novel explorations of endo-fucoidanases.
We compare the application of different modeling strategies in order to predict physical properties of five different industrial pectin formulations based on near‐infrared spectral data. Methods from the chemometric toolbox, such as partial least squares regression (PLS1 and PLS2) and ridge regression, were employed and compared to the performance of a 1‐D convolutional neural network (CNN). The pectin formulations were modeled in two major scenarios, individually using local models, and jointly using global models, which resulted in better prediction performance of the 1‐D CNN.
Remote sensing satellite images in the optical domain often contain missing or misleading data due to overcast conditions or sensor malfunctioning, concealing potentially important information. In this paper, we apply expectation maximization (EM) Tucker to NDVI satellite data from the Iberian Peninsula in order to gap-fill missing information. EM Tucker belongs to a family of tensor decomposition methods that are known to offer a number of interesting properties, including the ability to directly analyze data stored in multidimensional arrays and to explicitly exploit their multiway structure, which is lost when traditional spatial-, temporal- and spectral-based methods are used. In order to evaluate the gap-filling accuracy of EM Tucker for NDVI images, we used three data sets based on advanced very-high resolution radiometer (AVHRR) imagery over the Iberian Peninsula with artificially added missing data as well as a data set originating from the Iberian Peninsula with natural missing data. The performance of EM Tucker was compared to a simple mean imputation, a spatio-temporal hybrid method, and an iterative method based on principal component analysis (PCA). In comparison, imputation of the missing data using EM Tucker consistently yielded the most accurate results across the three simulated data sets, with levels of missing data ranging from 10 to 90%.
BACKGROUND:Culicoides biting midges transmit viruses resulting in disease in ruminants and equids such as bluetongue, Schmallenberg disease and African horse sickness. In the past decades, these diseases have led to important economic losses for farmers in Europe. Vector abundance is a key factor in determining the risk of vector-borne disease spread and it is, therefore, important to predict the abundance of Culicoides species involved in the transmission of these pathogens. The objectives of this study were to model and map the monthly abundances of Culicoides in Europe.METHODS:We obtained entomological data from 904 farms in nine European countries (Spain, France, Germany, Switzerland, Austria, Poland, Denmark, Sweden and Norway) from 2007 to 2013. Using environmental and climatic predictors from satellite imagery and the machine learning technique Random Forests, we predicted the monthly average abundance at a 1 km2 resolution. We used independent test sets for validation and to assess model performance.RESULTS:The predictive power of the resulting models varied according to month and the Culicoides species/ensembles predicted. Model performance was lower for winter months. Performance was higher for the Obsoletus ensemble, followed by the Pulicaris ensemble, while the model for Culicoides imicola showed a poor performance. Distribution and abundance patterns corresponded well with the known distributions in Europe. The Random Forests model approach was able to distinguish differences in abundance between countries but was not able to predict vector abundance at individual farm level.CONCLUSIONS:The models and maps presented here represent an initial attempt to capture large scale geographical and temporal variations in Culicoides abundance. The models are a first step towards producing abundance inputs for R0 modelling of Culicoides-borne infections at a continental scale.
Tick-borne pathogens cause diseases in animals and humans, and tick-borne disease incidence is increasing in many parts of the world. There is a need to assess the distribution of tick-borne pathogens and identify potential risk areas. We collected 29,440 tick nymphs from 50 sites in Scandinavia from August to September, 2016. We tested ticks in a real-time PCR chip, screening for 19 vector-associated pathogens. We analysed spatial patterns, mapped the prevalence of each pathogen and used machine learning algorithms and environmental variables to develop predictive prevalence models. All 50 sites had a pool prevalence of at least 33% for one or more pathogens, the most prevalent being Borrelia afzelii, B. garinii , Rickettsia helvetica , Anaplasma phagocytophilum, and Neoehrlichia mikurensis . There were large differences in pathogen prevalence between sites, but we identified only limited geographical clustering. The prevalence models performed poorly, with only models for R. helvetica and N. mikurensis having moderate predictive power (normalized RMSE from 0.74–0.75, R 2 from 0.43–0.48). The poor performance of the majority of our prevalence models suggest that the used environmental and climatic variables alone do not explain pathogen prevalence patterns in Scandinavia, although previously the same variables successfully predicted spatial patterns of ticks in the same area.
Ticks carry pathogens that can cause disease in both animals and humans, and there is a need to monitor the distribution and abundance of ticks and the pathogens they carry to pinpoint potential high risk areas for tick-borne disease transmission. In a joint Scandinavian study, we measured Ixodes ricinus instar abundance at 159 sites in southern Scandinavia in August-September, 2016, and collected 29,440 tick nymphs at 50 of these sites. We additionally measured abundance at 30 sites in August-September, 2017. We tested the 29,440 tick nymphs in pools of 10 in a Fluidigm real-time PCR chip to screen for 17 different tick-associated pathogens, 2 pathogen groups and 3 tick species. We present data on the geolocation, habitat type and instar abundance of the surveyed sites, as well as presence/absence of each pathogen in all analysed pools from the 50 collection sites and individual prevalence for each site. These data can be used alone or in combination with other data for predictive modelling and mapping of high-risk areas.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
During water stress, crops undertake adjustments in functional, structural, and biochemical traits. Hyperspectral data and machine learning techniques (PLS-R) can be used to assess water stress responses in plant physiology. In this study, we investigated the potential of hyperspectral optical (VNIR) measurements supplemented with thermal remote sensing and canopy height (hc) to detect changes in leaf physiology of soybean (C3) and maize (C4) plants under three levels of soil moisture in controlled environmental conditions. We measured canopy evapotranspiration (ET), leaf transpiration (Tr), leaf stomatal conductance (gs), leaf photosynthesis (A), leaf chlorophyll content and morphological properties (hc and LAI), as well as vegetation cover reflectance and radiometric temperature (TL,Rad). Our results showed that water stress caused significant ET decreases in both crops. This reduction was linked to tighter stomatal control for soybean plants, whereas LAI changes were the primary control on maize ET. Spectral vegetation indices (VIs) and TL,Rad were able to track these different responses to drought, but only after controlling for confounding changes in phenology. PLS-R modeling of gs, Tr, and A using hyperspectral data was more accurate when pooling data from both crops together rather than individually. Nonetheless, separated PLS-R crop models are useful to identify the most relevant variables in each crop such as TL,Rad for soybean and hc for maize under our experimental conditions. Interestingly, the most important spectral bands sensitive to drought, derived from PLS-R analysis, were not exactly centered at the same wavelengths of the studied VIs sensitive to drought, highlighting the benefit of having contiguous narrow spectral bands to predict leaf physiology and suggesting different wavelength combinations based on crop type. Our results are only a first but a promising step towards larger scale remote sensing applications (e.g., airborne and satellite). PLS-R estimates of leaf physiology could help to parameterize canopy level GPP or ET models and to identify different photosynthetic paths or the degree of stomatal closure in response to drought.
Recently, focus on tick-borne diseases has increased as ticks and their pathogens have become widespread and represent a health problem in Europe. Understanding the epidemiology of tickborne infections requires the ability to predict and map tick abundance. We measured Ixodes ricinus abundance at 159 sites in southern Scandinavia from August-September, 2016. We used field data and environmental variables to develop predictive abundance models using machine learning algorithms, and also tested these models on 2017 data. Larva and nymph abundance models had relatively high predictive power (normalized RMSE from 0.65-0.69, R-2 from 0.52-0.58) whereas adult tick models performed poorly (normalized RMSE from 0.94-0.96, R-2 from 0.04-0.10). Testing the models on 2017 data produced good results with normalized RMSE values from 0.59-1.13 and R-2 from 0.18-0.69. The resulting 2016 maps corresponded well with known tick abundance and distribution in Scandinavia. The models were highly influenced by temperature and vegetation, indicating that climate may be an important driver of I. ricinus distribution and abundance in Scandinavia. Despite varying results, the models predicted abundance in 2017 with high accuracy. The models are a first step towards environmentally driven tick abundance models that can assist in determining risk areas and interpreting human incidence data.
The taiga tick, Ixodes persulcatus, has previously been limited to eastern Europe and northern Asia, but recently its range has expanded to Finland and northern Sweden. The species is of medical importance, as it, along with a string of other pathogens, may carry the Siberian and Far Eastern subtypes of tick-borne encephalitis virus. These subtypes appear to cause more severe disease, with higher fatality rates than the central European subtype. Until recently, the meadow tick, Dermacentor reticulatus, has been absent from Scandinavia, but has now been detected in Denmark, Norway and Sweden. Dermacentor reticulatus carries, along with other pathogens, Babesia canis and Rickettsia raoultii. Babesia canis causes severe and often fatal canine babesiosis, and R. raoultii may cause disease in humans. We collected 600 tick nymphs from each of 50 randomly selected sites in Denmark, southern Norway and south-eastern Sweden in August-September 2016. We tested pools of 10 nymphs in a Fluidigm real time PCR chip to screen for I. persulcatus and D. reticulatus, as well as tick-borne pathogens. Of all the 30,000 nymphs tested, none were I. persulcatus or D. reticulatus. Our results suggest that I. persulcatus is still limited to the northern parts of Sweden, and have not expanded into southern parts of Scandinavia. According to literature reports and supported by our screening results, D. reticulatus may yet only be an occasional guest in Scandinavia without established populations.
Partial Least Squares (PLS) regression is a statistical method for supervised multivariate analysis. It relates two data blocks X and Y to each other with the aim of establishing a prediction model. When deployed in production, this model can be used to predict an outcome y from a newly measured feature vector x. PLS is popular in chemometrics, process control and other analytic fields, due to its striking advantages, namely the ability to analyze small sample sizes and the ability to handle high-dimensional data with cross-correlated features (where Ordinary Least Squares regression typically fails). In addition, and in contrast to many other machine learning approaches, PLS models can be interpreted using its latent variable structure just like principal components can be interpreted for a PCA analysis.
Efficient regeneration of NAD(P) + cofactors is essential for large‐scale application of alcohol dehydrogenases due to the high cost and chemical instability of these cofactors. NAD(P) + can be regenerated effectively using NAD(P)H oxidases (NOXs) that require molecular oxygen as a cosubstrate. In large‐scale biocatalytic processes, agitation and aeration are needed for sufficient oxygen transfer into the liquid phase, both of which have been shown to significantly increase the rate of enzyme deactivation. As such, the aim of this study was to identify the existence of a correlation between enzyme stability and gas–liquid interfacial area inside the bioreactor. This was done by measuring gas–liquid interfacial areas inside an aerated stirred reactor, using an in situ optical probe, and simultaneously measuring the kinetic stability of NOXs. Following enzyme incubation at various power inputs and gas‐phase compositions, the residual activity was assessed and video samples were analyzed through an image processing algorithm. Enzyme deactivation was found to be proportional to an increase in interfacial area up to a certain limit, where power input appears to have a higher impact. Furthermore, the presence of oxygen increased enzyme deactivation rates at low interfacial areas. The areas were validated with defined glass beads and found to be in the range of those in large‐scale bioreactors. Finally, a correlation between the enzyme half‐life and specific interfacial area was obtained. Therefore, we conclude that the method developed in this contribution can help to predict the behavior of biocatalyst stability under industrially relevant conditions, concerning specific gas–liquid interfacial areas.
Unlike satellite earth observation, multispectral images acquired by Unmanned Aerial Systems (UAS) provide great opportunities to monitor land surface conditions also in cloudy or overcast weather conditions. This is especially relevant for high latitudes where overcast and cloudy days are common. However, multispectral imagery acquired by miniaturized UAS sensors under such conditions tend to present low brightness and dynamic ranges, and high noise levels. Additionally, cloud shadows over space (within one image) and time (across images) are frequent in UAS imagery collected under variable irradiance and result in sensor radiance changes unrelated to the biophysical dynamics at the surface. To exploit the potential of UAS for vegetation mapping, this study proposes methods to obtain robust and repeatable reflectance time series under variable and low irradiance conditions. To improve sensor sensitivity to low irradiance, a radiometric pixel-wise calibration was conducted with a six-channel multispectral camera (mini-MCA6, Tetracam) using an integrating sphere simulating the varying low illumination typical of outdoor conditions at 55oN latitude. The sensor sensitivity was increased by using individual settings for independent channels, obtaining higher signal-to-noise ratios compared to the uniform setting for all image channels. To remove cloud shadows, a multivariate statistical procedure, Tucker tensor decomposition, was applied to reconstruct images using a four-way factorization scheme that takes advantage of spatial, spectral and temporal information simultaneously. The comparison between reconstructed (with Tucker) and original images showed an improvement in cloud shadow removal. Outdoor vicarious reflectance validation showed that with these methods, the multispectral imagery can provide reliable reflectance at sunny conditions with root mean square deviations of around 3%. The proposed methods could be useful for operational multispectral mapping with UAS under low and variable irradiance weather conditions as those prevalent in northern latitudes.
Background Tick-borne diseases have become increasingly common in recent decades and present a health problem in many parts of Europe. Control and prevention of these diseases require a better understanding of vector distribution. Aim Our aim was to create a model able to predict the distribution of Ixodes ricinus nymphs in southern Scandinavia and to assess how this relates to risk of human exposure. Methods We measured the presence of I. ricinus tick nymphs at 159 stratified random lowland forest and meadow sites in Denmark, Norway and Sweden by dragging 400 m transects from August to September 2016, representing a total distance of 63.6 km. Using climate and remote sensing environmental data and boosted regression tree modelling, we predicted the overall spatial distribution of I. ricinus nymphs in Scandinavia. To assess the potential public health impact, we combined the predicted tick distribution with human density maps to determine the proportion of people at risk. Results Our model predicted the spatial distribution of I. ricinus nymphs with a sensitivity of 91% and a specificity of 60%. Temperature was one of the main drivers in the model followed by vegetation cover. Nymphs were restricted to only 17.5% of the modelled area but, respectively, 73.5%, 67.1% and 78.8% of the human populations lived within 5 km of these areas in Denmark, Norway and Sweden. Conclusion The model suggests that increasing temperatures in the future may expand tick distribution geographically in northern Europe, but this may only affect a small additional proportion of the human population.
BACKGROUND:Biting midges of the genus Culicoides (Diptera: Ceratopogonidae) are small hematophagous insects responsible for the transmission of bluetongue virus, Schmallenberg virus and African horse sickness virus to wild and domestic ruminants and equids. Outbreaks of these viruses have caused economic damage within the European Union. The spatio-temporal distribution of biting midges is a key factor in identifying areas with the potential for disease spread. The aim of this study was to identify and map areas of neglectable adult activity for each month in an average year. Average monthly risk maps can be used as a tool when allocating resources for surveillance and control programs within Europe.METHODS:We modelled the occurrence of C. imicola and the Obsoletus and Pulicaris ensembles using existing entomological surveillance data from Spain, France, Germany, Switzerland, Austria, Denmark, Sweden, Norway and Poland. The monthly probability of each vector species and ensembles being present in Europe based on climatic and environmental input variables was estimated with the machine learning technique Random Forest. Subsequently, the monthly probability was classified into three classes: Absence, Presence and Uncertain status. These three classes are useful for mapping areas of no risk, areas of high-risk targeted for animal movement restrictions, and areas with an uncertain status that need active entomological surveillance to determine whether or not vectors are present.RESULTS:The distribution of Culicoides species ensembles were in agreement with their previously reported distribution in Europe. The Random Forest models were very accurate in predicting the probability of presence for C. imicola (mean AUC = 0.95), less accurate for the Obsoletus ensemble (mean AUC = 0.84), while the lowest accuracy was found for the Pulicaris ensemble (mean AUC = 0.71). The most important environmental variables in the models were related to temperature and precipitation for all three groups.CONCLUSIONS:The duration periods with low or null adult activity can be derived from the associated monthly distribution maps, and it was also possible to identify and map areas with uncertain predictions. In the absence of ongoing vector surveillance, these maps can be used by veterinary authorities to classify areas as likely vector-free or as likely risk areas from southern Spain to northern Sweden with acceptable precision. The maps can also focus costly entomological surveillance to seasons and areas where the predictions and vector-free status remain uncertain.
Laccases (EC 1.10.3.2) are enzymes known for their ability to catalyze the oxidation of phenolic compounds using molecular oxygen as the final electron acceptor. Laccase activity is commonly determined by monitoring spectrophotometric changes (absorbance) of the product or substrate during the enzymatic reaction. Fourier Transform Infrared Spectroscopy (FTIR) is a fast and versatile technique where spectral evolution profiling, i.e. assessment of the spectral changes of both substrate and products during enzymatic conversion in real time, can be used to assess enzymatic activity when combined with multivariate data analysis. We employed FTIR to monitor enzymatic oxidation of monolignols (sinapyl, coniferyl and p-coumaryl alcohol), sinapic acid, and sinapic aldehyde by four different laccases: three fungal laccases from Trametes versicolor, Trametes villosa and Ganoderma lucidum, respectively, and one bacterial laccase from Meiothermus ruber. By coupling the FTIR measurements with Parallel Factor Analysis (PARAFAC) we established a quantitative assay for assessing laccase activity. By combining PARAFAC modelling with Principal Component Analysis we show the usefulness of this technology as a multivariate tool able to compare and distinguish different laccase reaction patterns. We also demonstrate how the FTIR approach can be used to create a reference system for laccase activity comparison based on a relatively low number of measurements. Such a reference system has potential to function as a high-throughput method for comparing reaction pattern similarities and differences between laccases and hereby identify new and interesting enzyme candidates in large sampling pools.