Emerging infectious diseases demand comprehensive containment strategies, encompassing early detection and patient monitoring. Conventional diagnostic tests often suffer from low sensitivity, reliance on specialized equipment, and binary (positive/negative) outputs that provide limited clinical insight. To address these limitations, we introduce a novel approach integrating time-domain micro-nuclear magnetic resonance (TD-μNMR) with a dual intelligent relaxometric sensing (DIS) system, combining Fe3O4 and MnO nanoparticles with machine learning algorithms. This system profits from the distinct T1 and T2 relaxometric profiles of these nanoparticles to simultaneously detect endogenous (immune response) and exogenous (viral) analytes in complex biological samples. Specific molecular recognition by the nanoparticles induces measurable relaxometric changes, transforming them into sensitive NMR sensors. We demonstrate the efficacy of this method through multiplexed detection of SARS-CoV-2 antigens and antibodies, showcasing its potential for advanced diagnostics.
Accurate computer-aided cell detection in immunohistochemistry images of different tissues is essential for advancing digital pathology and enabling large-scale quantitative analysis. This paper presents a comprehensive comparison of six unsupervised segmentation methods against two supervised deep learning approaches for cell detection in immunohistochemistry images. The unsupervised methods are based on the continuity and similarity image properties, using techniques like clustering, active contours, graph cuts, superpixels, or edge detectors. The supervised techniques include the YOLO deep learning neural network and the U-Net architecture with heatmap-based localization for precise cell detection. All these methods were evaluated using leave-one-image-out cross-validation on the publicly available OIADB dataset, containing 40 oral tissue IHC images with over 40,000 manually annotated cells, assessed using precision, recall, and F1-score metrics. The U-Net model achieved the highest performance for cell nuclei detection, an F1-score of 75.3%, followed by YOLO with F1 = 74.0%, while the unsupervised OralImmunoAnalyser algorithm achieved only F1 = 46.4%. Although the two former are the best solutions for automatic pathological assessment in clinical environments, the latter could be useful for small research units without big computational resources.
Nutrient deficiency in wheat plants can lead to diseases and important losses in yield. These diseases can be visually detected on wheat leaf images. We perform the image classification as nutrient controlled or deficient, using a collection of 57 machine learning classifiers programmed in 4 different programming languages, applied on color texture features extracted from the images. We also use other 90 methods under the Caret automated R classification framework on the same features. Furthermore, we use 62 deep learning networks under three frameworks applied on the leaf images in three settings: trained from the scratch, fine-tuning of pretrained networks and classification of deep and shallow features extracted by deep networks. The radial basis function (RBF) neural network achieves the best performance, with kappa and accuracy of 57% and 81.2%, and with a low false positive rate (11.1%), while pretrained deep networks and classification of shallow features achieve 40% and 47%, respectively. Since nutrient deficiency is a continuous concept, ranging from 0% to 100%, and a sharp categorization into controlled and deficient may always be relative, these results identify the RBF network as an accurate approach for the detection of nutrient deficiency in wheat leaves.
Oral cancer ranks sixteenth amongst types of cancer by number of deaths. Many oral cancers are developed from potentially malignant disorders such as oral leukoplakia, whose most frequent predictor is the presence of epithelial dysplasia. Immunohistochemical staining using cell proliferation biomarkers such as ki67 is a complementary technique to improve the diagnosis and prognosis of oral leukoplakia. The cell counting of these images was traditionally done manually, which is time-consuming and not very reproducible due to intra- and inter-observer variability. The software presently available is not suitable for this task. This article presents the OralImmunoAnalyser software (registered by the University of Santiago de Compostela-USC), which combines automatic image processing with a friendly graphical user interface that allows investigators to oversee and easily correct the automatically recognized cells before quantification. OralImmunoAnalyser is able to count the number of cells in three staining levels and each epithelial layer. Operating in the daily work of the Odontology Faculty, it registered a sensitivity of 64.4% and specificity of 93% for automatic cell detection, with an accuracy of 79.8% for cell classification. Although expert supervision is needed before quantification, OIA reduces the expert analysis time by 56.5% compared to manual counting, avoiding mistakes because the user can check the cells counted. Hence, the SUS questionnaire reported a mean score of 80.9, which means that the system was perceived from good to excellent. OralImmunoAnalyser is accurate, trustworthy, and easy to use in daily practice in biomedical labs. The software, for Windows and Linux, with the images used in this study, can be downloaded from https://citius.usc.es/transferencia/software/oralimmunoanalyser for research purposes upon acceptance.
Breast cancer is the most diagnosed cancer worldwide and represents the fifth cause of cancer mortality globally. It is a highly heterogeneous disease, that comprises various molecular subtypes, often diagnosed by immunohistochemistry. This technique is widely employed in basic, translational and pathological anatomy research, where it can support the oncological diagnosis, therapeutic decisions and biomarker discovery. Nevertheless, its evaluation is often qualitative, raising the need for accurate quantitation methodologies. We present the software BreastAnalyser, a valuable and reliable tool to automatically measure the area of 3,3’-diaminobenzidine tetrahydrocholoride (DAB)-brown-stained proteins detected by immunohistochemistry. BreastAnalyser also automatically counts cell nuclei and classifies them according to their DAB-brown-staining level. This is performed using sophisticated segmentation algorithms that consider intrinsic image variability and save image normalization time. BreastAnalyser has a clean, friendly and intuitive interface that allows to supervise the quantitations performed by the user, to annotate images and to unify the experts’ criteria. BreastAnalyser was validated in representative human breast cancer immunohistochemistry images detecting various antigens. According to the automatic processing, the DAB-brown area was almost perfectly recognized, being the average difference between true and computer DAB-brown percentage lower than 0.7 points for all sets. The detection of nuclei allowed proper cell density relativization of the brown signal for comparison purposes between the different patients. BreastAnalyser obtained a score of 85.5 using the system usability scale questionnaire, which means that the tool is perceived as excellent by the experts. In the biomedical context, the connexin43 (Cx43) protein was found to be significantly downregulated in human core needle invasive breast cancer samples when compared to normal breast, with a trend to decrease as the subtype malignancy increased. Higher Cx43 protein levels were significantly associated to lower cancer recurrence risk in Oncotype DX-tested luminal B HER2- breast cancer tissues. BreastAnalyser and the annotated images are publically available https://citius.usc.es/transferencia/software/breastanalyser for research purposes.
The support vector machine (SVM) with Gaussian kernel often achieves state-of-the-art performance in classification problems, but requires the tuning of the kernel spread. Most optimization methods for spread tuning require training, being slow and not suited for large-scale datasets. We formulate an analytic expression to calculate, directly from data without iterative search, the spread minimizing the difference between Gaussian and ideal kernel matrices. The proposed direct gamma tuning (DGT) equals the performance of and is one to two orders of magnitude faster than the state-of-the art approaches on 30 small datasets. Combined with random sampling of training patterns, it also runs on large classification problems. Our method is very efficient in experiments with 20 large datasets up to 31 million of patterns, it is faster and performs significantly better than linear SVM, and it is also faster than iterative minimization.
Computer vision (CV) is a broad term mainly used to refer to processing image and video data [...]
Comorbidity between neurodevelopmental disorders is common, especially between autism spectrum disorder (ASD) and attention deficit/hyperactivity disorder (ADHD). This study aimed to detect overlapped sensory processing alterations in a sample of children and adolescents diagnosed with both ASD and ADHD. A collection of 42 standard and 8 proposed machine learning classifiers, 22 feature selection methods and 19 unbalanced classification strategies were applied on the 6 standard question groups of the Sensory Profile-2 questionnaire. The relatively low performance achieved by state-of-the-art classifiers led us to propose the feature population sum classifier, a probabilistic method based on class and feature value populations, designed for datasets where features are discrete numeric answers to questions in a questionnaire. The proposed method achieves the best kappa and accuracy, 60% and 82.5%, respectively, reaching 68% and 86.5% combined with backward sequential feature selection, with false positive and negative rates below 15%. Since the SP2 questionnaire can be filled by parents for children from three years, our prediction can alert the clinicians with an early diagnosis in order to apply early interventions.
Every day we live with more technology and computer systems. It is widely known that this technology is not neutral and can harm citizens, specially women. A possible cause for this lack of neutrality could be the under-representation of women on development teams. Our hypothesis is that the motivation of female students could increase through the use of alternative didactic methodologies, specifically the cooperative learning, including the gender dimension in its design. This study describes and analyses the use of cooperative learning to teach computer programming courses in a higher education STEM discipline, the Math degree. The statistical evaluation of the perception of the students along two academic years was very positive, specially for the female students.
The decision-making process for third molar removal or maintenance remains controversial in dental practice. The most important variables to be analyzed in predicting the potential of third molar eruption are retromolar space and the direction of eruption. The various methods for prediction include linear measures: measurement of the available space, mandibular size and growth, size of the third molar, and third molar angulation. The available software is not suitable for predicting third molar eruption. The purpose of the present work was to develop a clinical tool that can automatically predict eruption of the third molars based on combined linear and angular measurements. In this paper, the development and validation analysis of Panoramic Dental Application (PDApp) software (registered by the University of Santiago de Compostela (USC)) is presented, which can automatically predict third molar eruption from panoramic radiographs. This prediction is performed using a machine learning classifier (a support vector machine with Gaussian kernel) trained on a set of 188 cases wherein third molar angulation and the radiological retention coefficient are used as input data. Operating in the daily practice of the School of Dentistry at USC, an accuracy of 97.96% in predicting the potential of third molar eruption is achieved for a set of 539 third molars belonging to 289 patients. The software was also rated as the best imaginable system by the system usability scale (SUS) questionnaire. In this study, we developed and analyzed a new, unique software tool with increased diagnostic accuracy that will facilitate and optimize dental care in routine clinical workflow.
We propose the two-dimensional visual map classifier and regressor, which project the high-dimensional patterns on a 2D map, for human visualization and understanding of the data, and afterwards define a classification or regression map that predicts, for each 2D pattern, the class label (in classification) or the output value (in regression). The 2D projection is performed using the linear discriminant analysis, due to its high performance, speed and ability to project unseen (out-of-sample) patterns. The map is defined in an efficient way by assigning the proper output value to each square (or pixel) in the 2D map. The experiments show that the maps defined by both methods: (1) allow to understand visually the data distribution of a classification or regression problem; (2) their performances are very near to the state-of-the-art support vector classification and regression, including wrappers; and (3) they are very fast, between 1 and 5 orders of magnitude faster than the other approaches, spending less than 1 min to classify datasets with 5 million patterns. Matlab code is available.
Fish fecundity is one of the most relevant parameters for the estimation of the reproductive potential of fish stocks, used to assess the stock status to guarantee sustainable fisheries management. Fecundity is the number of matured eggs that each female fish can spawn each year. The stereological method is the most accurate technique to estimate fecundity using histological images of fish ovaries, in which matured oocytes must be measured and counted. A new segmentation technique, named the multi-scale Canny filter (MSCF), is proposed to recognize the boundaries of cells (oocytes), based on the Canny edge detector. Our results show the superior performance of MSCF on five fish species compared to five other state-of-the-art segmentation methods. It provides the highest F1 score in four out of five fish species, with values between 70% and 80%, and the highest percentage of correctly recognized cells, between 52% and 64%. This type of research aids in the promotion of sustainable fisheries management and conservation efforts, decreases research’s environmental impact and gives important insights into the health of fish populations and marine ecosystems.
Colorectal cancer (CRC) is one of the most common types of cancer worldwide. The KRAS mutation is present in 30–50% of CRC patients. This mutation confers resistance to treatment with anti-EGFR therapy. This article aims at proving that computer tomography (CT)-based radiomics can predict the KRAS mutation in CRC patients. The piece is a retrospective study with 56 CRC patients from the Hospital of Santiago de Compostela, Spain. All patients had a confirmatory pathological analysis of the KRAS status. Radiomics features were obtained using an abdominal contrast enhancement CT (CECT) before applying any treatments. We used several classifiers, including AdaBoost, neural network, decision tree, support vector machine, and random forest, to predict the presence or absence of KRAS mutation. The most reliable prediction was achieved using the AdaBoost ensemble on clinical patient data, with a kappa and accuracy of 53.7% and 76.8%, respectively. The sensitivity and specificity were 73.3% and 80.8%. Using texture descriptors, the best accuracy and kappa were 73.2% and 46%, respectively, with sensitivity and specificity of 76.7% and 69.2%, also showing a correlation between texture patterns on CT images and KRAS mutation. Radiomics could help manage CRC patients, and in the future, it could have a crucial role in diagnosing CRC patients ahead of invasive methods.
A simple and fast method, named ideal kernel tuning, is proposed to select the radial basis kernel spread of the support vector machine for classification. The spread is selected directly from data, without any training nor test, in order to bring the kernel matrix nearer to the ideal kernel matrix, with values one for patterns of the same class and zero otherwise. To avoid scaling with the training set size, the kernel matrix is calculated for a small set of class prototypes. The selected spread can be used also for multi-class classification problems considering the two most populated classes. Compared to other 5 popular tuning algorithms, the proposed approach is a smart and efficient strategy whose performance is very near to the state-of-the-art, is 2–4 orders of magnitude faster and requires very little memory, scales better with the dataset size both in time and memory, and is able to classify medium-size datasets up to 70,000 training patterns, where other methods fail.
Dry-cured ham is a traditional Mediterranean meat product consumed throughout the world. This product is very variable in terms of composition and quality. Consumer's acceptability of this product is influenced by different factors, in particular, visual intramuscular fat and its distribution across the slice, also known as marbling. On-line marbling assessment is of great interest for the industry for classification purposes. However, until now this assessment has been traditionally carried out by panels of experts and this methodology cannot be implement in industry. We propose a complete automatic system to predict marbling degree of dry-cured ham slices, which combines: (1) the color texture features of regions of interest (ROIs) extracted automatically for each muscle; and (2) machine learning models to predict the marbling. For the ROIs extraction algorithm more than the 90% of pixels of the ROI fall into the true muscle. The proposed system achieves a correlation of 0.92 using the support vector regression and a set of color texture features including statistics of each channel of RGB color image and Haralick's coefficients of its gray-level version. The mean absolute error was 0.46, which is lower than the standard desviation (0.5) of the marbling scores evaluated by experts. This high accuracy in the marbling prediction for sliced dry-cured ham would allow to deploy its application in the dry-cured ham industry.
The extreme learning machine (ELM) is a method to train single-layer feed-forward neural networks that became popular because it uses a fast closed-form expression for training that minimizes the training error with good generalization ability to new data. The ELM requires the tuning of the hidden layer size and the calculation of the pseudo-inverse of the hidden layer activation matrix for the whole training set. With large-scale classification problems, the computational overload caused by tuning becomes not affordable, and the activation matrix is extremely large, so the pseudo-inversion is very slow and eventually the matrix will not fit in memory. The quick extreme learning machine (QELM), proposed in the current paper, is able to manage large classification datasets because it: (1) avoids the tuning by using a bounded estimation of the hidden layer size from the data population; and (2) replaces the training patterns in the activation matrix by a reduced set of prototypes in order to avoid the storage and pseudo-inversion of large matrices. While ELM or even the linear SVM cannot be applied to large datasets, QELM can be executed on datasets up to 31 million data, 30,000 inputs and 131 classes, spending reasonable times (less than 1 h) in general purpose computers without special software nor hardware requirements and achieving performances similar to ELM.
The fish fecundity in an important parameter to manage sustainable fisheries. Traditionally, its calculation is manually performed on the histological images of fish ovaries, counting and measuring its matured reproductive cells, which is a very time consuming process. The automatization of this process implies the recognition of the matured cells in the image. This paper compares the statistical performance of five state-of-art segmentation techniques, which code is publically available, to segment histological images of fish ovary in order to recognize the outline of its cells. The approaches based on Canny filter and K-means clustering provides the best trade-off in performance and speed for both fish species tested. Although other approaches provide comparable performances at pixel level, we are interested in their efficiency at region level (the recognition of the outline of the cells) in order to measure them and calculate the fish fecundity.
The support vector machine (SVM) is a very important machine learning algorithm with state-of-the-art performance on many classification problems. However, on large datasets it is very slow and requires much memory. To solve this defficiency, we propose the fast support vector classifier (FSVC) that includes: 1) an efficient closed-form training free of any numerical iterative procedure; 2) a small collection of class prototypes that avoids to store in memory an excessive number of support vectors; and 3) a fast method that selects the spread of the radial basis function kernel directly from data, without classifier execution nor iterative hyper-parameter tuning. The memory requirements of FSVC are very low, spending in average only 6$\cdot 10^{-7}$·10-7 sec. per pattern, input and class, and processing datasets up to 31 millions of patterns, 30,000 inputs and 131 classes in less than 1.5 hours (less than 3 hours with only 2GB of RAM). In average, the FSVC is 10 times faster, requires 12 times less memory and achieves 4.7 percent more performance than Liblinear, that fails on the 4 largest datasets by lack of memory, being 100 times faster and achieving only 6.7 percent less performance than Libsvm. The time spent by FSVC only depends on the dataset size and thus it can be accurately estimated for new datasets, while Libsvm or Liblinear are much slower on “difficult” datasets, even if they are small. The FSVC adjusts its requirements to the available memory, classifying large datasets in computers with limited memory. Code for the proposed algorithm in the Octave scientific programming language is provided.1
Women are still underrepresented in science, technology, engineering, and mathematics (STEM) careers, not only in Spain but also in many Western and European countries. Female role models are key to broadening participation in STEM fields, which is why our teaching innovation group has been involved in initiatives to promote female vocations in STEM aimed at non-university education. In this paper we present different initiatives to provide female references to the students of the bachelor’s degree in Mathematics, Chemical Engineering, and master’s degree in computer vision. With these initiatives we aim to contribute to empowering female students of these degrees and to break with the sexist stereotypes that persist in our society. The activities have been well received by the students so we will continue with them in the next courses.
En este artículo analizamos la transversalización de la perspectiva de género en el campo de la Inteligencia Artificial (IA), cuyas aplicaciones influyen cada vez más en nuestras actividades cotidianas. Una disciplina altamente masculinizada, donde la mayoría de los profesionales son hombres y sus experiencias conforman y dominan la creación de algoritmos. Para reconocer la existencia de sesgos discriminatorios de género en los algoritmos y limitar sus consecuencias, es necesario introducir la perspectiva de género en estos estudios. En este documento revisamos el grado de introducción de competencias en género en los grados de IA en el estado español con el fin de mejorar la formación en género del alumnado.