The increasing reliance on Deep Learning models, combined with their inherent lack of transparency, has spurred the development of a novel field of study known as eXplainable AI (XAI) methods. These methods aim to enhance end-users' trust in automated systems by providing insights into the rationale behind their decisions. This paper presents a novel trust measure in XAI systems, allowing their refinement. Our proposed metric combines both performance metrics and trust indicators from an objective perspective. To validate this novel methodology, we conducted three case studies showing an improvement with respect to the state-of-the-art, with an increased sensitivity to different scenarios.
Background: The increasing use of telemedicine in surgical care has shown promise in improving patient outcomes and optimizing healthcare resources. Surgical site infections (SSIs) are a major cause of healthcare-associated infections (HAIs), leading to significant economic and health burdens. A pilot study already demonstrated that RedScar© achieved 100% sensitivity and 83.13% specificity in detecting SSIs. Patients reported high satisfaction regarding comfort, cost-effectiveness, and reduced absenteeism. Methods: This multicenter prospective study will include 168 patients undergoing abdominal surgery. RedScar© utilizes smartphone-based automated infection risk assessments without clinician input. App-based detection will be compared with in-person evaluations. Sensitivity and specificity will be analyzed using receiver operating characteristic (ROC) analysis, while secondary objectives include assessing patient satisfaction and standardizing telematic follow-up. Results: This study aims to evaluate the efficacy of the RedScar© app, sensitivity, specificity in detecting SSIs. Satisfaction regarding comfort, cost-effectiveness, and absenteeism due to telematic detection and the monitoring of SSIs will be recorded too. Conclusions: This study seeks to validate RedScar© as a reliable and scalable tool for postoperative monitoring. By improving early SSI detection, it has the potential to enhance surgical recovery, reduce healthcare costs, and optimize resource utilization.
Deep learning applied to chest X-ray (CXR) images has gained wide attention for its potential to improve diagnostic accuracy and accessibility in resource-limited healthcare settings. This study compares two deep learning strategies for lung disease classification: a Two-Stage approach that first detects abnormalities before classifying specific pathologies and a Direct multiclass classification approach. Using a curated database of CXR images covering diverse lung diseases, including COVID-19, pneumonia, pulmonary fibrosis, and tuberculosis, we evaluate the performance of various convolutional neural network architectures, the impact of lung segmentation, and explainability techniques. Our results show that the Two-Stage framework achieves higher diagnostic performance and fewer false positives than the Direct approach. Additionally, we highlight the limitations of segmentation and data augmentation techniques, emphasizing the need for further advancements in explainability and robust model design to support real-world diagnostic applications. Finally, we conduct a complementary evaluation of bone suppression techniques to assess their potential impact on disease classification performance.
Respiratory diseases remain the third leading cause of death in European Environment Agency member countries, with approximately 420,000 annual deaths and accounting for 6.1% of all EU deaths in 2021 [1,2]. Chest X-rays (CXR), due to their affordability and accessibility, are widely used in diagnostics and screening. However, their interpretation is often hindered by anatomical overlaps and variable pathology presentations, requiring expert radiologists whose availability may be limited in high-demand or resource-constrained environments. Deep learning offers a scalable solution to assist in diagnosis. This study presents a comparative evaluation of two deep learning strategies for CXR-based lung disease classification: (1) a Multiclass approach, in which a single model categorizes CXRs into five classes—four pathologies (pneumonia, COVID-19, tuberculosis, pulmonary fibrosis) and healthy cases; and (2) a Hierarchical approach, where an initial binary model detects abnormality, followed by a secondary model that classifies the specific pathology. This hierarchical strategy is intended to mimic clinical triage, prioritizing abnormality detection before disease differentiation. The inclusion of these four pathologies was motivated by their prevalence, radiographic ambiguity, and clinical relevance in triage settings. Attempts to further subclassify pneumonia into viral and bacterial types using CXRs alone proved unreliable, aligning with known clinical limitations [3] and supporting the selected disease scope. All models were based on DenseNet [4] architectures and trained using a curated, multi-institutional dataset combining over 20,000 images from 10 public and private sources. Images underwent preprocessing steps, including adaptive windowing, lung segmentation using a Weighted Ordered Weighted Averaging (WOWA) ensemble of six segmentation models [5], and targeted data augmentation. A consistent training setup using 5-fold cross-validation was used to ensure robust evaluation. To improve model transparency and clinical trust, we incorporated explainability methods (Grad-CAM [6], Grad-CAM++ [7], and Score-CAM [8]) and verified that activation maps focused on segmented lung regions. Coverage analysis confirmed that models trained with segmentation consistently concentrated their attention within anatomical boundaries, unlike unsegmented models that often relied on irrelevant cues. However, lesion-level validation using bounding boxes revealed limited spatial precision (IoU < 10%), highlighting ongoing challenges in XAI resolution. The Hierarchical approach achieved superior performance in terms of F1 scores for all pathological categories, including Pneumonia (0.9849 vs. 0.9771), COVID-19 (0.9078 vs. 0.9067), Pulmonary Fibrosis (0.9619 vs. 0.9516), and Tuberculosis (0.8926 vs. 0.8786) when compared to the multiclass classification strategy. Additionally, the F1 score for normal cases was also higher (0.9606 vs. 0.9365), indicating more reliable identification of healthy subjects. To evaluate statistical significance, a Shapiro-Wilk test confirmed the normality of the performance data (p > 0.05), enabling the use of a Students t-test. The results show that the Hierarchical Classification approach significantly outperformed the Multiclass Classification method in terms of overall classification performance (p = 0.0392), with a particularly significant improvement in the F1 score (p = 0.0196). A clinically critical metric, the False Positive Rate (FPR) for normal cases, was also notably reduced. The Hierarchical framework achieved a 33% lower FPR for normal cases compared to the multiclass model. This improvement is essential in medical applications, where incorrectly labeling diseased individuals as healthy can delay treatment and endanger patient safety. In summary, the proposed Hierarchical Classification framework demonstrates improved diagnostic performance and interpretability compared to a multiclass model. Its architecture, grounded in clinical logic and validated across multiple datasets, supports its potential for deployment in real-world environments lacking expert radiological oversight.
Sickle cell disease causes erythrocytes to become sickle-shaped, affecting their movement in the bloodstream and reducing oxygen delivery. It has a high global prevalence and places a significant burden on healthcare systems, especially in resource-limited regions. Automated classification of sickle cells in blood images is crucial, allowing the specialist to reduce the effort required and avoid errors when quantifying the deformed cells and assessing the severity of a crisis. Recent studies have proposed various erythrocyte representation and classification methods (Jennifer et al., 2023 [1]). Since classification depends solely on cell shape, a suitable approach models erythrocytes as closed planar curves in shape space (Epifanio et al., 2020). This approach employs elastic distances between shapes, which are invariant under rotations, translations, scaling, and reparameterizations, ensuring consistent distance measurements regardless of the curves’ position, starting point, or traversal speed. While previous methods exploiting shape space distances had achieved high accuracy, we refined the model by considering the geometric characteristics of healthy and sickled erythrocytes. Our method proposes (1) to employ a fixed parameterization based on the major axis of each cell to compute distances and (2) to align each cell with two templates using this parameterization before computing distances. Aligning shapes to templates before distance computation, a concept successfully applied in areas such as molecular dynamics (Richmond et al., 2004 [2]), and using a fixed parameterization, instead of minimizing distances across all possible parameterizations, simplifies calculations. This strategy achieves 96.03% accuracy rate in both supervised classification and unsupervised clustering. Our method ensures efficient erythrocyte classification, maintaining or improving accuracy over shape space models while significantly reducing computational costs.
The study of aggregation functions, either from a theoretical point of view or for their interesting applications, is a hot topic. In this paper we propose a new method for constructing aggregation functions on the set of Zadeh's discrete Z-numbers based on total orders. This method is characterised by using aggregation functions defined on the set of discrete fuzzy numbers whose support is a closed interval of the finite chain L-n = {0, 1, center dot center dot center dot, n}. Furthermore, it is shown that this construction method preserves important properties of the chosen initial aggregation functions.
Research on the construction of logical connectives using total (admissible) orders is a prolific area of study. Using such orders, a new method for constructing implication functions is defined on the set of discrete fuzzy numbers with support of a closed interval of a given finite chain and whose membership values belong to a finite set of fixed values. This method is based on the use of discrete implication functions defined on a finite chain. Furthermore, a bijective correspondence between the set of implication functions on the aforementioned subset of discrete fuzzy numbers and the set of discrete implication functions defined on the discrete chain is shown. Basic properties of these implication functions are thoroughly investigated, concluding that they are preserved under the proposed construction method. This result highlights the robustness and generality of the method, providing a systematic way to extend discrete implication functions to more complex structures while preserving their underlying properties.
This contribution presents a wavelet-based algorithm to detect patterns in images. A two-dimensional extension of the DST-II is introduced to construct adapted wavelets using the equation of the tensor product corresponding to the diagonal coefficients in the 2D discrete wavelet transform. A 1D filter was then estimated that meets finite energy conditions, vanished moments, orthogonality, and four new detection conditions. These allow, when performing the 2D transform, for the filter to detect the pattern by taking the diagonal coefficients with values of the normalized similarity measure, defined by Guido, as greater than 0.7, and α=0.1. The positions of these coefficients are used to estimate the position of the pattern in the original image. This strategy has been used successfully to detect artificial patterns and localize mass-like abnormalities in digital mammography images. In the case of the latter, high sensitivity and positive predictive value in detection were achieved but not high specificity or negative predictive value, contrary to what occurred in the 1D strategy. This means that the proposed detection algorithm presents a high number of false negatives, which can be explained by the complexity of detection in these types of images.
El uso de wavelets adaptadas para el reconocimiento de patrones es muy atractivo por la multiescalaridad de la transformada wavelet. Sin embargo, el buen desempeño de estos algoritmos en la detección de patrones depende fuertemente de la construcción de los filtros adaptados al patrón de interés. La Transformada Shapelet Discreta II [9] (DST-II) es un algoritmo inspirado en la transformada wavelet, que permite el diseño de filtros a la medida para la detección de patrones en señales unidimensionales. La construcción de estos filtros requiere la solución de un sistema de ecuaciones no lineales, que según [9] se puede efectuar mediante cualquier método iterativo. Esta investigación presenta un novedoso y exhaustivo estudio numérico que demuestra el impacto de la elección del método numérico adecuado para la solución del sistema no lineal en la DST-II. La eficacia de los filtros estimados repercute en el desempeño de esta transformada en la detección de patrones. Los mejores resultados se obtienen al combinar el método de Newton con preiteración mediante el algoritmo de continuación. La convergencia alcanzada para el 55, 37% de los patrones sugiere que la DST-II podría ser adecuada para patrones con formas específicas, de utilidad en aplicaciones sobre señales biomédicas.
Automatic segmentation of organs and regions in medical imaging is a valuable tool for specialists. This study explores various automatic methods for lung segmentation in X-ray images. First, various neural network architectures are applied for this segmentation task, and subsequently, they are ranked based on their performance through statistical analysis. Then, as some architectures are more suitable for segmenting certain structures or regions of the X-ray, aggregation and consensus methods are studied to fuse the various neural network segmentations, with the aim of obtaining a more complete segmentation. The study reveals that the method based on the WOWA aggregation function, coupled with a maximum-based consensus method, statistically outperforms the individual segmentation provided by the best-performing neural network.
This study explores aggregation and consensus methods to combine lung segmentations from various neural network models in X-ray images, aiming to enhance accuracy and completeness. Through extensive experimentation, the research identifies the most effective aggregation method, with WOWA aggregation and a maximum-based consensus approach outperforming individual models. This underscores the importance of aggregation techniques in optimizing anatomical structure segmentation in medical imaging.
Background/Objectives: This study assessed the feasibility and security of remote surgical wound monitoring using the RedScar© smartphone app, which employs automated diagnosis for early visual detection of infections without direct healthcare personnel involvement. Additionally, patient satisfaction with telematic care was evaluated as a secondary aim. Surgical site infection (SSI) is the second leading cause of healthcare-associated infections (HAIs), leading to prolonged hospital stays, heightened patient distress, and increased healthcare costs. Methods: The study employed a prospective paired-cohort and single-blinded design, with a sample size of 47 adult patients undergoing abdominal surgery. RedScar© was used for remote telematic monitoring, evaluating the feasibility and security of this approach. A satisfaction questionnaire assessed patient experience. The study protocol was registered at ClinicalTrials.gov under the identifier NCT05485233. Results: Out of 47 patients, 41 successfully completed both remote and in-person follow-ups. RedScar© demonstrated a sensitivity of 100% in detecting SSIs, with a specificity of 83.13%. The kappa coefficient of 0.8171 indicated substantial agreement between the application’s results and human observers. Patient satisfaction with telemonitoring was high: 97.6% believed telemonitoring reduces costs, 90.47% perceived it prevents work/school absenteeism, and 80.9% found telemonitoring comfortable. Conclusions: This is the first study to evaluate an automatic smartphone application on real patients for diagnosing postoperative wound infections. It establishes the safety and feasibility of telematic follow-up using the RedScar© application for surgical wound assessment. The high sensitivity suggests its utility in identifying true cases of infection, highlighting its potential role in clinical practice. Future studies are needed to address limitations and validate the efficacy of RedScar© in diverse patient populations.
This study presents a method to improve state-of-the-art concave point detection methods as the first step towards effectively segmenting overlapping objects in images. The approach relies on analysing the curvature of the object contour. This method comprises three main steps. First, the original image is preprocessed to obtain the curvature value at each contour point. Second, the regions with higher curvatures are selected and a recursive algorithm is applied to refine previously selected regions. Finally, a concave point is obtained for each region by analysing the relative position of their neighbourhood. Furthermore, the experimental results indicate that improving the detection of concave points leads to better division of clusters. To evaluate the quality of the concave point detection algorithm, a synthetic dataset was constructed to simulate the presence of overlapping objects. This dataset includes the precise location of concave points, which serve as the ground truth for evaluation. As a case study, the performance of a well-known application, such as the splitting of overlapping cells in images of peripheral blood smears samples from patients with sickle cell anaemia, was evaluated. We used the proposed method to detect concave points in cell clusters and then separated these clusters by ellipse fitting.
Symmetry is one of the distinguishing features when diagnosing the malignancy of skin lesions. In this work, we introduce an extension of the SymDerm dataset with around 2000 new annotations, and analyze 1) the effect of different data augmentation techniques on learning the skin lesion symmetry classification task, and 2) how the learning of this task is affected when combined with the classification of its malignancy in a multitask learning environment. We conclude that, although not all data augmentation techniques improve classification performance, these techniques achieve an increase of approximately 7.7% for B.Acc and Precision, 8.0% for Recall and F1-score, and 15.08% for the Kappa score. Moreover, we show that symmetry classification benefits from the introduction of an auxiliary task by stabilizing the learning curve and decreasing the train-validation learning gap.
Deep learning techniques provide a powerful and versatile tool in different areas, such as object segmentation in medical images. In this paper, we propose a network based on the U-Net architecture to perform the segmentation of wounds and staples in abdominal surgery images. Moreover, since both tasks are highly interdependent, we propose a multitask architecture that allows to simultaneously obtain, in the same network evaluation, the masks with the staples and wound location of the image. When performing this multitasking, it is necessary to formulate a global loss function that linearly combines the losses of both partial tasks. This is why the study also involves the GradNorm algorithm to determine which weight is associated to each loss function during each training step. The main conclusion of the study is that multitask segmentation offers superior performance compared to segmenting by separate tasks.
Image noise can be viewed as unwanted disturbances in a digital image that should be removed or reduced before further processing and analysis. Impulsive noise, also known as impulse noise, is a very disruptive type of noise, characterized by abrupt variations in brightness in a subset of the image pixels. Impulsive noise commonly occurs during image acquisition and transmission. To mitigate its effects, various impulsive noise reduction methods have been proposed by the image processing community. In contrast to classical filters such as the median filter, most current impulsive noise reduction techniques implement a two-step approach that consists of a noise detection phase to identify noisy pixels and a filtering phase to reduce the amount of noise in the presumably corrupted pixels. The approach presented in this paper is also along this line. To be more precise, we draw on the principles of two state-of-the-art impulsive noise reduction methods, namely the adaptive fuzzy transform based image filter (ATIF) and the improved fuzzy mathematical morphology open-close filter (i-FMMOCS), in order to propose a new method for general impulsive noise reduction.
Skin cancer has become a public health problem due to its increasing incidence. However, the malignancy risk of the lesions can be reduced if diagnosed at an early stage. To do so, it is essential to identify particular characteristics such as the symmetry of lesions. In this work, we present a novel approach for skin lesion symmetry classification of dermoscopic images based on deep learning techniques. We use a CNN model, which classifies the symmetry of a skin lesion as either "fully asymmetric", "symmetric with respect to one axis", or "symmetric with respect to two axes". Moreover, we introduce a new dataset of labels for 615 skin lesions. During the experimentation framework, we also evaluate whether it is beneficial to rely on transfer learning from pre-trained CNNs or traditional learning-based methods. As a result, we present a new simple, robust and fast classification pipeline that outperforms methods based on traditional approaches or pre-trained networks, with a weighted-average F1-score of 64.5%.
This paper presents a method that uses a sequential representation to train Hidden Markov Models as an algorithm for the supervised morphological classification of erythrocytes in peripheral blood samples from patients with sickle cell anemia, considering three classes: circular, elongated and with others deformations. This sequential learning method provides the probability of belonging the object to the class and for the representation of the red cell contour, characteristics are not obtained, but the contour is analyzed as a sequence of curvatures. The experimentation carried out analyzes each group as a class and considers 3, 8, 9, 10 and 11 states, so that the method is capable of dealing with the local angular differences existing in this representation, with the aim of improving the performance of the classification obtained so far. To check the effectiveness of this method, we use samples with balanced classes formed by images of individual erythrocytes, in similar amounts for each of the three classes. Measurements of sensitivity, precision, specificity, F1 and classification accuracy were obtained. The best results were obtained for the representation considering 10 states.
Multitasking learning improves a model’s ability to generalize by learning multiple tasks in parallel. However, it is difficult to know how each task influences the others’ learning. In this work, we study in-depth the behavior of the tasks of skin lesion segmentation, hair mask segmentation, and the inpainting of those hairs, in a multitasking framework to discover how they influence each other. The experiments are performed using an encoder-decoder convolutional neural network and images from five public databases: PH2, dermquest, dermis, EDRA2002, and the ISIC Data Archive. To evaluate the tasks’ performance, we use a series of metrics on which we apply a statistical test to check the superiority of each task in a multitasking model with respect to their individual performance. We also check, in a three-task model, whether there is a task that dominates the learning stage. Finally, we conclude that while the inpainting task does not benefit from this type of learning, the rest of the tasks improve their performance when compared to that obtained by their corresponding single-task model.