Microscopy images often suffer from high levels of noise, which can hinder further analysis and interpretation. Content-aware image restoration (CARE) methods have been proposed to address this issue, but they often require large amounts of training data and suffer from over-fitting. To overcome these challenges, we propose a novel framework for few-shot microscopy image denoising. Our approach combines a generative adversarial network (GAN) trained via contrastive learning (CL) with two structure preserving loss terms (Structural Similarity Index and Total Variation loss) to further improve the quality of the denoised images using little data. We demonstrate the effectiveness of our method on three well-known microscopy imaging datasets, and show that we can drastically reduce the amount of training data while retaining the quality of the denoising, thus alleviating the burden of acquiring paired data and enabling few-shot learning. The proposed framework can be easily extended to other image restoration tasks and has the potential to significantly advance the field of microscopy image analysis.
Nowadays, there is still a challenge in virtual reality to obtain an accurate displacement prediction of the user. This could be a future key element to apply in the so-called redirected walking methods. Meanwhile, deep learning provides us with new tools to reach greater achievements in this type of prediction. Specifically, long short-term memory recurrent neural networks obtained promising results recently. This gives us clues to continue researching in this line to predict virtual reality user’s displacement. This manuscript focuses on the collection of positional data and a subsequent new way to train a deep learning model to obtain more accurate predictions. The data were collected with 44 participants and it has been analyzed with different existing prediction algorithms. The best results were obtained with a new idea, the use of rotation quaternions and the three dimensions to train the previously existing models. The authors strongly believe that there is still much room for improvement in this research area by means of the usage of new deep learning models.
This paper investigates the ability of deep neural networks (DNNs) to predict a specific room impulse response (RIR) between two given room locations. We use three end-to-end deep learning (DL) models based on the encoder-decoder structure: an auto-encoder (AE), a variational AE (VAE) and a UNet. They try to generate a new RIR given: 1) the short-time Fourier transform (STFT) of a true RIR from the same room, 2) the spatial coordinates of the true and new RIRs, and 3) several room-related parameters. On the one hand, the magnitude and phase of the STFT are computed and presented to the DNN input. On the other hand, the spatial coordinates and the room parameters form an information vector that is embedded into the DNNs through their latent space (AE, VAE) or equivalent layer (UNet). A real database of RIRs measured in five different rooms was used to train and test the models. Two experiments were carried out to study the influence of the magnitude and phase terms of the loss function on the performance of the models. An additional experiment investigated the ability of the DL models to generalize across rooms. Results show that the three DNNs are able to predict the magnitude of the STFT, but they cannot accurately predict its phase. When comparing their ability to generalize across spaces, the UNet achieves the best results.
The creation of artistic images through the use of Artificial Intelligence is an area that has been gaining interest in recent years. In particular, the ability of Neural Networks to separate and subsequently recombine the style of different images, generating a new artistic image with the desired style, has been a source of study and attraction for the academic and industrial community. This work addresses the challenge of generating artistic images that are framed in the style of pictorial Impressionism and, specifically, that imitate the style of one of its greatest exponents, the painter Claude Monet. After having analysed several theoretical approaches, the Cycle Generative Adversarial Networks are chosen as base model. From this point, a new training methodology which has not been applied to cyclical systems so far, the top-k approach, is implemented. The proposed system is characterised by using in each iteration of the training those k images that, in the previous iteration, have been able to better imitate the artist’s style. To evaluate the performance of the proposed methods, the results obtained with both methodologies, basic and top-k, have been analysed from both a quantitative and qualitative perspective. Both evaluation methods demonstrate that the proposed top-k approach recreates the author’s style in a more successful manner and, at the same time, also demonstrate the ability of Artificial Intelligence to generate something as creative as impressionist paintings.
Gunshot detection in natural environments is crucial for the protection of endangered species. In this work, we present a novel dataset built from the soundscape recording at five different locations of the Spanish Albufera National Park. We then carry out an experimental study to detect gunshots from the rest of the background sounds and noises labeled as “vbackground”. For this purpose, we perform comprehensive experiments both in the data input and the modeling stages of three efficient deep convolutional neural networks (DCNNs), obtaining an F1 score (harmonic mean of precision and recall) of 0.92 for the best model. The best three DCNNs are also used to monitor one hour of the Albufera soundscape where the gunshot class represents 8 % of the testset. The recall values obtained with our model are comparable to previous works monitoring gunshots in real scenarios.
Image restoration is an important task in a wide variety of topics. Especially in the domain of microscopy images, several content-aware image restoration (CARE) methods have arisen to improve the interpretability of acquired data. One of the main problems is the presence of high levels of noise that must be removed before any further post-processing or analysis happens. To solve this issue, we propose a simple and yet effective framework consisting of a generative adversarial network (GAN) coupled with a regularization term (differentiable data augmentation) that highly increases the quality of the denoising for three different well-known microscopy imaging data sets. We also introduce two structure preserving loss terms (Structural Similarity Index and Total Variation loss) that, added to our framework, help to further improve the quality of the results. In addition, we prove that our method is able to generalize well when trained on one dataset and used in another. Finally, we show that we can drastically reduce the amount of training data while retaining the quality of the denoising, thus alleviating the burden of acquiring paired data and enabling few-shot learning.
Facial information is processed by our brain in such a way that we immediately make judgments about, for example, attractiveness or masculinity or interpret personality traits or moods of other people. The appearance of each facial feature has an effect on our perception of facial traits. This research addresses the problem of measuring the size of these effects for five facial features (eyes, eyebrows, nose, mouth, and jaw). Our proposal is a mixed feature-based and image-based approach that allows judgments to be made on complete real faces in the categorization tasks, more than on synthetic, noisy, or partial faces that can influence the assessment. Each facial feature of the faces is automatically classified considering their global appearance using principal component analysis. Using this procedure, we establish a reduced set of relevant specific attributes (each one describing a complete facial feature) to characterize faces. In this way, a more direct link can be established between perceived facial traits and what people intuitively consider an eye, an eyebrow, a nose, a mouth, or a jaw. A set of 92 male faces were classified using this procedure, and the results were related to their scores in 15 perceived facial traits. We show that the relevant features greatly depend on what we are trying to judge. Globally, the eyes have the greatest effect. However, other facial features are more relevant for some judgments like the mouth for happiness and femininity or the nose for dominance.
In 2015 we began a sub-challenge at the EndoVis workshop at MICCAI in Munich using endoscope images of ex-vivo tissue with automatically generated annotations from robot forward kinematics and instrument CAD models. However, the limited background variation and simple motion rendered the dataset uninformative in learning about which techniques would be suitable for segmentation in real surgery. In 2017, at the same workshop in Quebec we introduced the robotic instrument segmentation dataset with 10 teams participating in the challenge to perform binary, articulating parts and type segmentation of da Vinci instruments. This challenge included realistic instrument motion and more complex porcine tissue as background and was widely addressed with modifications on U-Nets and other popular CNN architectures. In 2018 we added to the complexity by introducing a set of anatomical objects and medical devices to the segmented classes. To avoid over-complicating the challenge, we continued with porcine data which is dramatically simpler than human tissue due to the lack of fatty tissue occluding many organs.
Classification or typology systems used to categorize different human body parts have existed for many years. Nevertheless, there are very few taxonomies of facial features. Ergonomics, forensic anthropology, crime prevention or new human-machine interaction systems and online activities, like e-commerce, e-learning, games, dating or social networks, are fields in which classifications of facial features are useful, for example, to create digital interlocutors that optimize the interactions between human and machines. However, classifying isolated facial features is difficult for human observers. Previous works reported low inter-observer and intra-observer agreement in the evaluation of facial features. This work presents a computer-based procedure to automatically classify facial features based on their global appearance. This procedure deals with the difficulties associated with classifying features using judgements from human observers, and facilitates the development of taxonomies of facial features. Taxonomies obtained through this procedure are presented for eyes, mouths and noses.
We present a different approach for annotating laparoscopic images for segmentation in a weak fashion and experimentally prove that its accuracy when trained with partial cross-entropy is close to that obtained with fully supervised approaches.
This work describes a new hybrid method for accurate iris segmentation from full-face images independently of the ethnicity of the subject. It is based on a combination of three methods: facial key-point detection, integro-differential operator (IDO) and mathematical morphology. First, facial landmarks are extracted by means of the Chehra algorithm in order to obtain the eye location. Then, the IDO is applied to the extracted sub-image containing only the eye in order to locate the iris. Once the iris is located, a series of mathematical morphological operations is performed in order to accurately segment it. Results are obtained and compared among four different ethnicities (Asian, Black, Latino and White) as well as with two other iris segmentation algorithms. In addition, robustness against rotation, blurring and noise is also assessed. Our method obtains state-of-the-art performance and shows itself robust with small amounts of blur, noise and/or rotation. Furthermore, it is fast, accurate, and its code is publicly available.
Optical coherence tomography (OCT) is a useful technique to monitor retinal damage. We present an automatic method to accurately classify rodent OCT images in healthy and pathological (before and after 14 days of intravitreal injection of Endothelin-1, respectively) making use of the DenseNet-201 architecture fine-tuned and a customized top-model. We validated the performance of the method on 1912 OCT images yielding promising results (\(AUC = 0.99\pm 0.01\) in a \(P=15\) leave-P-out cross-validation). Besides, we also compared the results of the fine-tuned network with those achieved training the network from scratch, obtaining some interesting insights. The presented method poses a step forward in understanding pathological rodent OCT retinal images, as at the moment there is no known discriminating characteristic which allows classifying this type of images accurately. The result of this work is a very accurate and robust automatic method to distinguish between healthy and a rodent model of glaucoma, which is the backbone of future works dealing with human OCT images.
Deep Neural Networks (DNN) have become a powerful, and extremely popular mechanism, which has been widely used to solve problems of varied complexity, due to their ability to make models fitted to non-linear complex problems. Despite its well-known benefits, DNNs are complex learning models whose parametrisationand architecture are made usually by hand. This paper proposes a new Evolutionary Algorithm, named EvoDeep, devoted to evolve the parameters and the architecture of a DNN in order to maximise its classification accuracy, as well as maintaining a valid sequence of layers. This model is tested against a widely used dataset of handwritten digits images. The experiments performed using this dataset show that the Evolutionary Algorithm is able to select the parameters and the DNN architecture appropriately, achieving a 98.93% accuracy in the best run.
We are constantly making very fast attributions from faces, such as whether a person is trustworthy or threatening, that influence our behavior towards people. In this work, we present a method to automatically tell the importance of facial features on social trait perception. We employ an unsupervised clustering method to group the facial features by similarity and then create a model which explains the contribution of each facial feature to each social trait by means of a Genetic Algorithm. Our model deals with the difficulties associated to quantifying social impression using judgments from human observers (low inter- and intra-observer agreement) and obtains significant correlations greater than 0.7 for all social impressions, which justifies the method developed. Finally, the weights of the eyebrows, eyes, nose, mouth, jawline and facial feature distances are shown and discussed. This work poses a step forward in social trait impression understanding, as to the date, there is no other work quantifying the effects of facial features on social trait perception.
Human faces play a central role in our lives. Thanks to our behavioural capacity to perceive faces, how a face looks in a painting, a movie, or an advertisement can dramatically influence what we feel about them and what emotions are elicited. Facial information is processed by our brain in such a way that we immediately make judgements like attractiveness or masculinity or interpret personality traits or moods of other people. Due to the importance of appearance-driven judgements of faces, this has become a major focus not only for psychological research, but for neuroscientists, artists, engineers, and software developers. New technologies are now able to create realistic looking synthetic faces that are used in arts, online activities, advertisement, or movies. However, there is not a method to generate virtual faces that convey the desired sensations to the observers. In this work, we present a genetic algorithm based procedure to create realistic faces combining facial features in the adequate relative positions. A model of how observers will perceive a face based on its features’ appearances and relative positions was developed and used as the fitness function of the algorithm. The model is able to predict 15 facial social traits related to aesthetic, moods, and personality. The proposed procedure was validated comparing its results with the opinion of human observers. This procedure is useful not only for creating characters with artistic purposes, but also for online activities, advertising, surgery, or criminology.
Classification or typology systems used to categorize different human body parts exist for many years. Nevertheless, there are very few taxonomies of facial features. A reason for this might be that classifying isolated facial features is difficult for human observers. Previous works reported low inter-observer and intra-observer agreement in the evaluation of facial features. Therefore, this work presents a computer-based procedure to classify facial features based on their global appearance automatically. First, facial features are located, extracted and aligned using a facial landmark detector. Then, images are characterized using the eigenpictures approach. We then perform a clustering of each type of facial feature using as input the weights extracted from the eigenpictures approach. Finally, we validate the obtained clusterings with humans. This procedure deals with the difficulties associated with classifying features using judgments from human observers and facilitates the development of taxonomies of facial features. Taxonomies obtained with this procedure are presented for eyes, noses and mouths.
Deep Neural Networks (DNN) have become a powerful, widely used, and successful mechanism to solve problems of different nature and varied complexity. Their ability to build models adapted to complex non-linear problems, have made them a technique widely applied and studied. One of the fields where this technique is currently being applied is in the malware classification problem. The malware classification problem has an increasing complexity, due to the growing number of features needed to represent the behaviour of the application as exhaustively as possible. Although other classification methods, as those based on SVM, have been traditionally used, the DNN pose a promising tool in this field. However, the parameters and architecture setting of these DNNs present a serious restriction, due to the necessary time to find the most appropriate configuration. This paper proposes a new genetic algorithm designed to evolve the parameters, and the architecture, of a DNN with the goal of maximising the malware classification accuracy, and minimizing the complexity of the model. This model is tested against a dataset of malware samples, which are represented using a set of static features, so the DNN has been trained to perform a static malware classification task. The experiments carried out using this dataset show that the genetic algorithm is able to select the parameters and the DNN architecture settings, achieving a 91% accuracy.
The purpose of the present study is to investigate whether the effectiveness of a new ad on digital channels (YouTube) can be predicted by using neural networks and neuroscience-based metrics (brain response, heart rate variability and eye tracking). Neurophysiological records from 35 participants were exposed to 8 relevant TV Super Bowl commercials. Correlations between neurophysiological-based metrics, ad recall, ad liking, the ACE metrix score and the number of views on YouTube during a year were investigated. Our findings suggest a significant correlation between neuroscience metrics and self-reported of ad effectiveness and the direct number of views on the YouTube channel. In addition, and using an artificial neural network based on neuroscience metrics, the model classifies (82.9% of average accuracy) and estimate the number of online views (mean error of 0.199). The results highlight the validity of neuromarketing-based techniques for predicting the success of advertising responses. Practitioners can consider the proposed methodology at the design stages of advertising content, thus enhancing advertising effectiveness. The study pioneers the use of neurophysiological methods in predicting advertising success in a digital context. This is the first article that has examined whether these measures could actually be used for predicting views for advertising on YouTube.
This work focuses on finding the most discriminatory or representative features that allow to classify commercials according to negative, neutral and positive effectiveness based on the Ace Score index. For this purpose, an experiment involving forty-seven participants was carried out. In this experiment electroencephalography (EEG), electrocardiography (ECG), Galvanic Skin Response (GSR) and respiration data were acquired while subjects were watching a 30-min audiovisual content. This content was composed by a submarine documentary and nine commercials (one of them the ad under evaluation). After the signal pre-processing, four sets of features were extracted from the physiological signals using different state-of-the-art metrics. These features computed in time and frequency domains are the inputs to several basic and advanced classifiers. An average of 89.76% of the instances was correctly classified according to the Ace Score index. The best results were obtained by a classifier consisting of a combination between AdaBoost and Random Forest with automatic selection of features. The selected features were those extracted from GSR and HRV signals. These results are promising in the audiovisual content evaluation field by means of physiological signal processing.