Satellite imagery provides society with data to solve a wide range of problems, including monitoring environmental changes or the detection of objects of interest, among many others. Anomaly detection has gained attention in recent years, as it can be used to address the aforementioned problems effectively. Traditional anomaly detection methods pose challenges, as they rely on labelled datasets, which are often scarce, expensive to obtain, and limited in capturing the full spectrum of real-world anomalies. To address these challenges, self-supervised learning has emerged as a powerful paradigm that leverages the usage of anomaly detection techniques on unlabelled data, allowing for cheaper solutions that can learn robust representations of known and unknown anomalies in many fields, including earth observation-related fields. In this study, we investigate state-of-the-art self-supervised anomaly detection methods and assess their performance on two widely used satellite imagery datasets: LandCover.ai, which focuses on land cover classification, and the High-Resolution Cloud Detection Dataset, which provides detailed cloud cover annotations. These datasets enable a comprehensive evaluation of anomaly detection effectiveness across diverse geospatial contexts. Nevertheless, existing state-of-the-art models often exhibit limitations, particularly in achieving a balanced trade-off between precision and recall, largely due to the scarcity of supervisory information inherent to the anomaly detection task. To address this challenge, we propose a novel ensemble self-supervised learning framework that integrates both low- and high-level abstraction objectives, thereby leveraging local and global feature representations to improve anomaly detection performance. Experimental results demonstrate that our approach consistently outperforms existing self-supervised methods, achieving improvements in F1 score ranging from 2 to 10
BACKGROUND:The development of artificial intelligence (AI) models for medical image analysis remains a significant challenge. Main issues include a lack of model generalization, the need for extensive datasets for training, and poor model robustness for clinical application. Current evaluation frameworks often fail to provide subject-based insights necessary to identify model biases and address domain shifts. METHODS:To address these limitations, this study introduces AUDIT, an open-source Python library designed to enhance the evaluation of segmentation models and the analysis of MRI datasets. AUDIT includes modules for extracting region-specific features and calculating performance metrics, as well as a dynamic web APP for interactive model evaluation and data exploration. RESULTS:All the source code, tutorials, and documentation can be found at https://github.com/caumente/AUDIT, and can directly be installed from the Python Package Index using pip install auditapp. Through a series of use cases commonly encountered in AI-driven brain tumor segmentation, this manuscript highlights the versatility and broad applicability of the AUDIT library across diverse scenarios. CONCLUSIONS:AUDIT was designed to fill the existing gaps in the literature on the evaluation of AI segmentation models, ultimately aiming to advance the AI medical image analysis domain. Beyond its extensive analysis capabilities, AUDIT has been designed to support integration with external libraries and applications.
Recommender systems are widely used to help users find relevant and personalized items in various domains. However, providing accurate recommendations is not enough to ensure user satisfaction, trust, and engagement. Nowadays, users demand transparency from these systems typically in the form of an explanation of the recommendation given. This paper presents a novel explainable Recommender System designed to generate recommendations from natural language queries while providing model-intrinsic explanations inspired by attention mechanisms. The system adds transparency, interpretability, and new user cold-start capabilities. We evaluate our approach on twelve datasets from diverse domains and languages, demonstrating its effectiveness and robustness. Results show that our proposal achieves competitive accuracy with respect to strong baselines, while consistently outperforming a prior interpretable model developed for the same task.
The properties and performance of biodiesels depend on the composition of the fatty acid methyl esters (FAMEs) present. Obtaining new biodiesels is neither a cheap nor a simple task. Therefore, it is important to model their properties in advance in order to know what composition would be needed to fulfill a given requirement. This paper presents a multitask approach that models the behavior of biodiesels using preference learning and also obtains representations of those biodiesels in a feature space by means of an autoencoder. The performance of the system in different configurations is also analyzed, showing its robustness. A case study was conducted to analyze the cetane number, which is an important property in biodiesel. Correlations between the cetane number and the distribution of FAMEs have been identified. These correlations are more clearly observed when the FAMEs are grouped according to common characteristics, such as the percentage of saturated fatty acids (SFA), monounsaturated fatty acids (MUFA), and polyunsaturated fatty acids (PUFA) in the biodiesel, or the average number of double bonds (DB). Based on these groupings, it can be observed that the cetane number decreases with increasing unsaturation. Additionally, it was determined that the cetane number value has a minimal or no correlation with the average molecular weight (Mw) of the biodiesel.
Background: Ultrasound (US) is a medical imaging modality that plays a crucial role in the early detection of breast cancer. The emergence of numerous deep learning systems has offered promising avenues for the segmentation and classification of breast cancer tumors in US images. However, challenges such as the absence of data standardization, the exclusion of non-tumor images during training, and the narrow view of single-task methodologies have hindered the practical applicability of these systems, often resulting in biased outcomes. This study aims to explore the potential of multi-task systems in enhancing the detection of breast cancer lesions. Methods: To address these limitations, our research introduces an end-to-end multi-task framework designed to leverage the inherent correlations between breast cancer lesion classification and segmentation tasks. Additionally, a comprehensive analysis of a widely utilized public breast cancer ultrasound dataset named BUSI was carried out, identifying its irregularities and devising an algorithm tailored for detecting duplicated images in it. Results: Experiments are conducted utilizing the curated dataset to minimize potential biases in outcomes. Our multi-task framework exhibits superior performance in breast cancer respecting single-task approaches, achieving improvements close to 15% in segmentation and classification. Moreover, a comparative analysis against the state-of-the-art reveals statistically significant enhancements across both tasks. Conclusion: The experimental findings underscore the efficacy of multi-task techniques, showcasing better generalization capabilities when considering all image types: benign, malignant, and non-tumor images. Consequently, our methodology represents an advance towards more general architectures with real clinical applications in the breast cancer field.
Few-shot learning is crucial for downstream tasks involving point clouds, given the challenge of obtaining sufficient datasets due to extensive collecting and labeling efforts. Pre-trained VLM-Guided point cloud models, containing abundant knowledge, can compensate for the scarcity of training data, potentially leading to very good performance. However, adapting these pre-trained point cloud models to specific few-shot learning tasks is challenging due to their huge number of parameters and high computational cost. To this end, we propose a novel Dynamic Multimodal Prompt Tuning method, named DMMPT, for boosting few-shot learning with pre-trained VLM-Guided point cloud models. Specifically, we build a dynamic knowledge collector capable of gathering task- and data-related information from various modalities. Then, a multimodal prompt generator is constructed to integrate collected dynamic knowledge and generate multimodal prompts, which efficiently direct pre-trained VLM-guided point cloud models toward few-shot learning tasks and address the issue of limited training data. Our method is evaluated on benchmark datasets not only in a standard N-way K-shot few-shot learning setting, but also in a more challenging setting with all classes and K-shot few-shot learning. Notably, our method outperforms other prompt-tuning techniques, achieving highly competitive results comparable to full fine-tuning methods while significantly enhancing computational efficiency.
In their first year of university, students perceive some courses as least related to their degree, and hence, show a lack of motivation in them, dedicating more time to those that they consider as more related to their degree. This leads to a decrease in students’ performance in these courses, even when they have sufficient capacity to do well. Moreover, it is observed that students are not aware of their knowledge level regarding the course content and generally overestimate it. This paper presents a method to increase students’ engagement in parts of the courses that are most difficult. This method was tested in one course, and the results support its effectiveness.
Navigation through large volumes of images is a complex and tedious task that requires tools to facilitate the exploration and discovery of visual information. Photo summaries are one of these tools, which consist of selecting a reduced set of images that best represent the original data source. However, creating photo summaries in the context of recommender systems poses several challenges: How to select the most relevant images for each item? How to encode each image? How to evaluate the quality of the generated summary? In this manuscript, we propose a clustering-based method to create a visual summary in the context of a restaurant recommender system, which includes the photos taken by users who visited the restaurants (items) in a given city. These photos are encoded using a deep neural network that takes into account not only their content but also the relationships between users and restaurants. This encoding will allow us to create a visual summary that captures the essence of user tastes and illustrates the gastronomic offer of the city. We also propose a similarity measure between items based on the users who have visited them and an evaluation method that calculates to what extent the summary obtained represents the original data source. The experimentation carried out includes five datasets and the obtained results demonstrate the adequacy of our proposal for the construction of these summaries.
Recommender Systems (RS) are based on the generalization of the observed interactions of a population of users with a collection of items. Collaborative Filters (CF) give good results, but they degrade when there are few interactions to learn from. The alternative would be to observe some features of the users that could be linked to their tastes. However, specific information on users or items is often not available. In this research work, we explore how to exploit the photos of items taken by users. Our aim is to assign similar meanings to the photos of items with which the same group of users interacted. For this purpose, we define a multi-label classification task from images to sets of users. The classifier uses a general-purpose convolutional neural network to extract the basic visual features, followed by additional layers necessary to accomplish the learning task. To evaluate our proposal we compared it with CFs, using two tourism datasets that include: restaurants of six cities and points of interest of three locations. According to the experimentation carried out, the poor results achieved by CFs are outperformed by our proposal, which takes into account the visual and taste semantics of the available photos.
Studying microRNA (miRNAs) in certain agri-food products is attractive because (1) they have potential as biomarkers that may allow traceability and authentication of such products; and (2) they may reveal insights into the products' functional potential. The present study evaluated differences in miRNAs levels in fat and cellular fractions of tank milk collected from commercial farms which employ extensive or intensive dairy production systems. We first sequenced miRNAs in three milk samples from each production system, and then validated miRNAs whose levels in the cellular and fat fraction differed significantly between the two production systems. To accomplish this, we used quantitative PCR with both fractions of tank milk samples from another 20 commercial farms. Differences in miRNAs were identified in fat fractions: overall levels of miRNAs, and, specifically, the levels of bta-mir-215, were higher in intensive systems than in extensive systems. Bovine mRNA targets for bta-miR-215 and their pathway analysis were performed. While the causes of these miRNAs differences remain to be elucidated, our results suggest that the type of production system could affect miRNAs levels and potential functionality of agri-food products of animal origin.
Recommender Systems are a very useful tool which let companies and service providers focus in the preferences of their customers, helping them to avoid an overwhelming variety of choices. In this context, clustering tools can play an important role to detect groups of customers with similar tastes. Thus, companies can make personalized marketing campaigns, offering to their users new products which have been consumed by other users with comparable preferences. In this paper we present a general framework to cluster users with respect to their tastes when the registers stored about the interactions between users and products are extremely scarce. Commonly, clustering methods employ the values of features describing the samples to be clustered (users in our case), but such features are not always available. We propose some alternative representations for users, in which their tastes are gathered to some extent, so that clustering algorithms can take advantage and make more homogeneous groups in this regard. To illustrate the performance of the whole framework, we tested it on six popular datasets commonly used as a benchmark for recommender systems, as well as on an extremely sparse real-world dataset that records the preferences of readers to click promoted links in digital publications. In the experimental section we compare our proposed representations to other common user encodings. We show that clustering users attending only to their feature values or to the items they have evaluated gives rise to the worst scores in terms of taste homogeneity.
Abstract Recommender systems have proven their usefulness both for companies and customers. The former increase their sales and the latter get a more satisfying shopping experience. These systems can benefit from the advent of explainable artificial intelligence, since a well-explained recommendation will be more convincing and may broaden the customer’s purchasing options. Many approaches offer justifications for their recommendations based on the similarity (in some sense) between users, past purchases, etc., which require some knowledge of the users. In this paper we present a recommender system with explanatory capabilities which is able to deal with the so-called cold-start problem, since it does not require any previous knowledge of the user. Our method learns the relationship between the products and some relevant words appearing in the textual reviews written by previous customers for those products. Then, starting from the textual query of a user’s request for recommendation, our approach elaborates a list of products and explains each recommendation on the basis of the compatibility between the query’s words and the relevant terms for each product.
En los primeros cursos de los grados universitarios se encuentran las asignaturas que los alumnos perciben como las menos relacionadas con el grado que estudian. Algunos alumnos, ante esta situación, muestran una falta de motivación en esas asignaturas dedicando más tiempo y dedicación a asignaturas que ellos consideran más afines a su grado. Esto provoca una disminución en el rendimiento de los alumnos en estas asignaturas aun cuando tienen capacidades suficientes para cursar de manera provechosa dichas asignaturas. Además, se observa en estas asignaturas que los alumnos no son verdadermante conscientes de su nivel de conocimiento de la asignatura, normalmente sobreestimándolo. En esta comunicación, se analiza el caso de una asignatura que se encuentra en tal situación, se propone un procedimiento para tratar de solucionar este problema y se aplica en uno de los grupos de teoría de dicha asignatura. Este procedimiento incluye autoevaluación por parte de los alumnos y evaluación por pares. Los resultados obtenidos muestran la efectividad del procedimiento propuesto.
The aim of Recommender Systems is to suggest items (products) to satisfy each user’s particular taste. Representation strategies play a very important role in these systems, as an adequate codification of users and items is expected to ease the induction of a model which synthesizes their tastes and make better recommendations. However, in addition to gathering information about users’ tastes, there is an additional aspect that can be relevant for a proper codification strategy, namely the order in which the user interacted with the items. In this paper, several encoding strategies based on neural networks are analyzed and applied to solve two different recommendation tasks in the context of music playlists. The results show that the order in which the musical pieces were listened to is relevant for the codification of items (songs). We also find that the encoding of user profiles should use a different amount of historical data depending on the learning task to be solved. In other words, we do not always have to use all the available data; sometimes, it is better to discard old information, as tastes change over time.
Explaining the output of a complex system, such as a Recommender System (RS), is becoming of utmost importance for both users and companies. In this paper we explore the idea that personalized explanations can be learned as recommendation themselves. There are plenty of online services where users can upload some photos, in addition to rating items. We assume that users take these photos to reinforce or justify their opinions about the items. For this reason we try to predict what photo a user would take of an item, because that image is the argument that can best convince her of the qualities of the item. In this sense, an RS can explain its results and, therefore, increase its reliability. Furthermore, once we have a model to predict attractive images for users, we can estimate their distribution. Thus, the companies acquire a vivid knowledge about the aspects that the clients highlight of their products. The paper includes a formal framework that estimates the authorship probability for a given pair (user, photo). To illustrate the proposal, we use data gathered from TripAdvisor containing the reviews (with photos) of restaurants in six cities of different sizes.
The study of marine plankton data is vital to monitor the health of the world's oceans. In recent decades, automatic plankton recognition systems have proved useful to address the vast amount of data collected by specially engineered in situ digital imaging systems. At the beginning, these systems were developed and put into operation using traditional automatic classification techniques, which were fed with hand-designed local image descriptors (such as Fourier features), obtaining quite successful results. In the past few years, there have been many advances in the computer vision community with the rebirth of neural networks. In this paper, we leverage how descriptors computed using convolutional neural networks trained with out-of-domain data are useful to replace hand-designed descriptors in the task of estimating the prevalence of each plankton class in a water sample. To achieve this goal, we have designed a broad set of experiments that show how effective these deep features are when working in combination with state-of-the-art quantification algorithms.
The articles in the long tail are those that are not popular in some sense, but all together often represent a large proportion of the products covered by a recommender system. For companies, it is important to recommend these items that otherwise could be unknown to their customers. It is also interesting for users because knowing about these items might constitute a pleasant surprise. But long-tail items are not the only we might wish to recommend. Thus, some companies promote products on seasonal offers. It is a challenge to manage the preferences on items whose interaction with users is scarce. There is a trade-off between recommending items that users like and those belonging to a certain kind. We present a framework to address recommendations where the items will have a weight that quantifies our interest in recommending them in a broad sense. Then, we derive a factorization method that optimizes the award of the recommendations. To test the method, we present an exhaustive experimentation with a real-world dataset on digital news. We show that it is possible to improve dramatically the novelty (those items of special interest) and diversity of items with a tiny penalization in the accuracy.
Label noise and class imbalance are two of the critical challenges when training image-based deep neural networks, especially in the biomedical image processing domain. Our work focuses on how to address the two challenges effectively and accurately in the task of lesion segmentation from biomedical/medical images. To address the pixel-level label noise problem, we propose an advanced transfer training and learning approach with a detailed DICOM pre-processing method. To address the tumor/non-tumor class imbalance problem, we exploit a self-adaptive fully convolutional neural network with an automated weight distribution mechanism to spot the Radiomics lung tumor regions accurately. Furthermore, an improved conditional random field method is employed to obtain sophisticated lung tumor contour delineation and segmentation. Finally, our approach has been evaluated using several well-known evaluation metrics on the Lung Tumor segmentation dataset used in the 2018 IEEE VIP-CUP Challenge. Experimental results show that our weakly supervised learning algorithm outperforms other deep models and state-of-the-art approaches.
Oscar Luaces合作论文数Artificial Intelligence Center, University of Oviedo at Gijon39