Integrating deep learning into healthcare enables personalized care but raises trust issues due to model opacity. To improve transparency, we propose a system for dental age estimation from panoramic images that combines an opaque and a transparent method within a Natural Language Generation module. This module generates clinician-friendly textual explanations of age estimations, designed with dental experts through a rule-based approach. Following the best practices in the field, the quality of the generated explanations was manually validated by human experts using a questionnaire. The results showed a strong performance, since the experts rated 4.77±0.12 (out of 5) on average across the five dimensions considered. We also performed a trustworthy self-assessment procedure following the Assessment List for Trustworthy Artificial Intelligence checklist, in which it scored 4.40±0.27 (out of 5) across its seven dimensions.
Cardiovascular analysis based on cardiac imaging highly benefits from automated segmentation methods. The use of deep learning is currently the most effective approach. However, its utilization in medical settings is frequently constrained by the unavailability of high-capacity hardware resources, and the high variability of medical images, which challenges the generalization ability of deep learning techniques. We propose a pipeline of two sequential U-Net for CT and MRI segmentation, configured with low complexity, allowing for usability in clinical practice. In the first stage, single-label segmentation is used to crop the image volume to a bounding box surrounding the heart. The second stage focuses on the detected region of interest. Multi-label segmentation is performed on the trimmed volume to extract 7 different substructures of the heart. Results from the WHS++ challenge validation phase show that our method achieves an average Dice Similarity Coefficient of 0.9311 on CT, and 0.8652 on MRI data. Importantly, the inference times are kept to a minimum, even when using CPU computing ( ∼ 7 s).
In a given species, genomes and 16S rRNA gene sequences, along with their intragenomic copy numbers, can vary greatly across environments. The gene copy numbers are crucial for technologies which estimate microbial abundances based on gene counts, such as polymerase chain reaction and high-throughput sequencing. In these, taxa with fewer genes may be underestimated, while those with more genes might be overestimated. Therefore, it is essential to have accurate gene copy number databases specific to the niche under study. The 16S rRNA Gene Oral Sequences dataset (16SGOSeq) contains the number of 16S rRNA genes and their variants in the complete genomes of the bacterial and archaeal species present in the human oral cavity. It includes 3,192 complete genomes of oral bacteria and 191 complete genomes of oral archaea, from which the 16S rRNA gene sequences were extracted, and the sequence variants were identified. This oral-specific dataset of prokaryotic organisms and the pipeline followed for its construction can be applied by clinical microbiologists, bioinformaticians, or microbial ecologists in future microbiome research.
Being born small carries significant health risks, including increased neonatal mortality and a higher likelihood of future cardiac diseases. Accurate estimation of gestational age is critical for monitoring fetal growth, but traditional methods, such as estimation based on the last menstrual period, are in some situations difficult to obtain. While ultrasound-based approaches offer greater reliability, they rely on manual measurements that introduce variability. This study presents an interpretable deep learning-based method for automated gestational age calculation, leveraging a novel segmentation architecture and distance maps to overcome dataset limitations and the scarcity of segmentation masks. Our approach achieves performance comparable to state-of-the-art models while reducing complexity, making it particularly suitable for resource-constrained settings and with limited annotated data. Furthermore, our results demonstrate that the use of distance maps is particularly suitable for estimating femur endpoints.
One area of bioinformatics that is currently attracting particular interest is the classification of polymicrobial diseases using machine learning (ML), with data obtained from high-throughput amplicon sequencing of the 16S rRNA gene in human microbiome samples. The microbial dysbiosis underlying these types of diseases is particularly challenging to classify, as the data is highly dimensional, with potentially hundreds or even thousands of predictive features. In addition, the imbalance in the composition of the microbial community is highly heterogeneous across samples. In this paper, we propose a curated pipeline for binary phenotype classification based on a count table of 16S rRNA gene amplicons, which can be applied to any microbiome. To evaluate our proposal, raw 16S rRNA gene sequences from samples of healthy and periodontally affected oral microbiomes that met certain quality criteria were downloaded from public repositories. In the end, a total of 2,581 samples were analysed. In our approach, we first reduced the dimensionality of the data using feature selection methods. After tuning and evaluating different machine learning (ML) models and ensembles created using Dynamic Ensemble Selection (DES) techniques, we found that all DES models performed similarly and were more robust than individual models. Although the margin over other methods was minimal, DES-P achieved the highest AUC and was therefore selected as the representative technique in our analysis. When diagnosing periodontal disease with saliva samples, it achieved with only 13 features an F1 score of 0.913, a precision of 0.881, a recall (sensitivity) of 0.947, an accuracy of 0.929, and an AUC of 0.973. In addition, we used EPheClass to diagnose inflammatory bowel disease (IBD) and obtained better results than other works in the literature using the same dataset. We also evaluated its effectiveness in detecting antibiotic exposure, where it again demonstrated competitive results. This highlights the importance and generalisation aspect of our classification approach, which is applicable to different phenotypes, study niches, and sample types. The code is available at https://gitlab.citius.usc.es/lara.vazquez/epheclass.
Abstract Background No clinical trials have evaluated the antimicrobial activity and substantivity of gel formulations containing chlorhexidine (CHX) and cymenol. Objective To compare the in situ antimicrobial effect and substantivity of a new 0.20% CHX + cymenol gel (test) with the current 0.20% CHX gel formulation (control) on salivary flora and dental plaque biofilm up to seven hours after a single application. Methods A randomised-crossover clinical trial was conducted with 29 orally healthy volunteers participating in the development of Experiments 1 (saliva) and 2 (dental plaque biofilm). All subjects participated in both experiments and were randomly assigned to receive either the test or control gels. Samples were collected at baseline and five minutes and one, three, five, and seven hours after a single application of the products. The specimens were processed using confocal laser scanning microscopy after staining with the LIVE/DEAD® BacLight™ solution. Bacterial viability (BV) was quantified in the saliva and biofilm samples. The BV was calculated using the DenTiUS Biofilm software. Results In Experiment 1, the mean baseline BV was significantly reduced five minutes after application in the test group (87.00% vs. 26.50%; p < 0.01). This effect was maintained throughout all sampling times and continued up to seven hours (40.40%, p < 0.01). The CHX control followed the same pattern. In Experiment 2, the mean baseline BV was also significantly lower five minutes after applying the test gel for: (1) the total thickness of biofilm (91.00% vs. 5.80%; p < 0.01); (2) the upper layer (91.29% vs. 3.94%; p < 0.01); and (3) the lower layer (86.29% vs. 3.83%; p < 0.01). The reduction of BV from baseline was observed for the full-thickness and by layers at all sampling moments and continued seven hours after application (21.30%, 24.13%, and 22.06%, respectively; p < 0.01). Again, the control group showed similar results. No significant differences between test and control gels were observed in either saliva or dental plaque biofilm at any sampling time. Conclusions A 0.20% CHX + cymenol gel application demonstrates potent and immediate antimicrobial activity on salivary flora and de novo biofilm. This effect is maintained seven hours after application. Similar effects are obtained with a 0.20% CHX-only gel.
BACKGROUND:The selection of primer pairs in sequencing-based research can greatly influence the results, highlighting the need for a tool capable of analysing their performance in-silico prior to the sequencing process. We therefore propose PrimerEvalPy, a Python-based package designed to test the performance of any primer or primer pair against any sequencing database. The package calculates a coverage metric and returns the amplicon sequences found, along with information such as their average start and end positions. It also allows the analysis of coverage for different taxonomic levels.RESULTS:As a case study, PrimerEvalPy was used to test the most commonly used primers in the literature against two oral 16S rRNA gene databases containing bacteria and archaea. The results showed that the most commonly used primer pairs in the oral cavity did not match those with the highest coverage. The best performing primer pairs were found for the detection of oral bacteria and archaea.CONCLUSIONS:This demonstrates the importance of a coverage analysis tool such as PrimerEvalPy to find the best primer pairs for specific niches. The software is available under the MIT licence at https://gitlab.citius.usc.es/lara.vazquez/PrimerEvalPy .
Background The effect of cymenol mouthwashes on levels of dental plaque has not been evaluated thus far. Objective To analyse the short-term, in situ, anti-plaque effect of a 0.1% cymenol mouthwash using the DenTiUS Deep Plaque software. Methods Fifty orally healthy participants were distributed randomly into two groups: 24 received a cymenol mouthwash for eight days (test group A) and 26 a placebo mouthwash for four days and a cymenol mouthwash for a further four days thereafter (test group B). They were instructed not to perform other oral hygiene measures. On days 0, 4, and 8 of the experiment, a rinsing protocol for staining the dental plaque with sodium fluorescein was performed. Three intraoral photographs were taken per subject under ultraviolet light. The 504 images were analysed using the DenTiUS Deep Plaque software, and visible and total plaque indices were calculated (ClinicalTrials ID NCT05521230). Results On day 4, the percentage area of visible plaque was significantly lower in test group A than in test group B (absolute = 35.31 ± 14.93% vs. 46.57 ± 18.92%, p = 0.023; relative = 29.80 ± 13.97% vs. 40.53 ± 18.48%, p = 0.024). In comparison with the placebo, the cymenol mouthwash was found to have reduced the growth rate of the area of visible plaque in the first four days by 26% (absolute) to 28% (relative). On day 8, the percentage areas of both the visible and total plaque were significantly lower in test group A than in test group B (visible absolute = 44.79 ± 15.77% vs. 65.12 ± 16.37%, p < 0.001; visible relative = 39.27 ± 14.33% vs. 59.24 ± 16.90%, p < 0.001; total = 65.17 ± 9.73% vs. 74.52 ± 13.55%, p = 0.007). Accounting for the growth rate with the placebo mouthwash on day 4, the above results imply that the cymenol mouthwash in the last four days of the trial reduced the growth rate of the area of visible plaque (absolute and relative) by 53% (test group A) and 29% (test group B), and of the area of total plaque by 48% (test group A) and 41% (test group B). Conclusions The 0.1% cymenol mouthwash has a short-term anti-plaque effect in situ, strongly conditioning the rate of plaque growth, even in clinical situations with high levels of dental plaque accumulation.
Sequencing has been widely used to study the composition of the oral microbiome present in various health conditions. The extent of the coverage of the 16S rRNA gene primers employed for this purpose has not, however, been evaluated in silico using oral-specific databases. This paper analyses these primers using two databases containing 16S rRNA sequences from bacteria and archaea found in the human mouth and describes some of the best primers for each domain. A total of 369 distinct individual primers were identified from sequencing studies of the oral microbiome and other ecosystems. These were evaluated against a database reported in the literature of 16S rRNA sequences obtained from oral bacteria, which was modified by our group, and a self-created oral archaea database. Both databases contained the genomic variants detected for each included species. Primers were evaluated at the variant and species levels, and those with a species coverage (SC) ≥75.00
Developing robust and performant methods for diagnosing COVID-19, particularly for triaging processes, is crucial. This study introduces a completely automated system to detect COVID-19 by means of the analysis of Chest X-Ray scans (CXR). The proposed methodology is based on few-shot techniques, enabling to work on small image datasets. Moreover, a set of additions have been done to enhance the diagnostic capabilities. First, a network to extract the lung region to rely only on the most relevant image area. Second, a new cost function to penalize each misclassification according to the clinical consequences. Third, a system to combine different predictions from the same image to increase the robustness of the diagnoses. The proposed approach was validated on the public dataset COVIDGR-1.0, yielding a classification accuracy of 79.10% ± 3.41% and, thus, outperforming other state-of-the-art methods. In conclusion, the proposed methodology has proven to be suitable for the diagnosis of COVID-19.
Dental radiographies have been used for many decades for estimating the chronological age, with a view to forensic identification, migration flow control, or assessment of dental development, among others. This study aims to analyse the current application of chronological age estimation methods from dental X-ray images in the last 6 years, involving a search for works in the Scopus and PubMed databases. Exclusion criteria were applied to discard off-topic studies and experiments which are not compliant with a minimum quality standard. The studies were grouped according to the applied methodology, the estimation target, and the age cohort used to evaluate the estimation performance. A set of performance metrics was used to ensure good comparability between the different proposed methodologies. A total of 613 unique studies were retrieved, of which 286 were selected according to the inclusion criteria. Notable tendencies to overestimation and underestimation were observed in some manual approaches for numeric age estimation, being especially notable in the case of Demirjian (overestimation) and Cameriere (underestimation). On the other hand, the automatic approaches based on deep learning techniques are scarcer, with only 17 studies published in this regard, but they showed a more balanced behaviour, with no tendency to overestimation or underestimation. From the analysis of the results, it can be concluded that traditional methods have been evaluated in a wide variety of population samples, ensuring good applicability in different ethnicities. On the other hand, fully automated methods were a turning point in terms of performance, cost, and adaptability to new populations.
Hundreds of publications have studied the oral microbiome through 16S rRNA gene sequencing. However, none have assessed the number of 16S rRNA genes in the genomes of oral microbes, or how the use of primer pairs targeting different regions affects the detection of MAs from different taxa.
In the past few years, one area of bioinformatics that has sparked special interest is the classification of diseases using machine learning. This is especially challenging in solving the classification of dysbiosis-based diseases, i.e., diseases caused by an imbalance in the composition of the microbial community. In this work, a curated pipeline is followed for classifying phenotypes using 16S rRNA gene amplicons, focusing on Crohn’s disease. It aims to reduce the dimensionality of data through a feature selection step, decreasing the computational cost, and maintaining an acceptably high f1-score. From this study, an ensemble model is proposed to contain the best-performing techniques from several representative machine learning algorithms. High f1-scores of up to 0.81 were reached thanks to this ensemble joining multilayer perceptron, extreme gradient boosting, and support vector machines, with as low as 300 target number of features. The results achieved were similar to or even better than other works studying the same data, so we demonstrated the goodness of our method.
The accuracy obtained with deep learning-based systems usually depends on the availability of large image datasets, which is not always possible. Consequently, it is necessary to apply techniques to increase the size of these datasets and their variability in a reliable and efficient way. In this respect, a novel tool for image augmentation and the associated data, IDALib, is presented. It provides an automatic method to perform transformation operations jointly on the images and related data, such as landmarks or masks, with the main aim of minimising the computational overhead. Thus, it applies automatically a set of optimisations, such as the vectorisation and composition of operations. Furthermore, the transformations are performed in GPU, which leads to a notable speedup, specially in dual GPU setups. IDALib is publicly available in the PyPI repository, https://pypi.org/project/ida-lib/
Although clustering by operational taxonomic units (OTUs) is widely used in the oral microbial literature, no research has specifically evaluated the extent of the limitations of this sequence clustering-based method in the oral microbiome. Consequently, our objectives were to: 1) evaluate in-silico the coverage of a set of previously selected primer pairs to detect oral species having 16S rRNA sequence segments with ≥97% similarity; 2) describe oral species with highly similar sequence segments and determine whether they belong to distinct genera or other higher taxonomic ranks. Thirty-nine primer pairs were employed to obtain the in-silico amplicons from the complete genomes of 186 bacterial and 135 archaeal species. Each fasta file for the same primer pair was inserted as subject and query in BLASTN for obtaining the similarity percentage between amplicons belonging to different oral species. Amplicons with 100% alignment coverage of the query sequences and with an amplicon similarity value ≥97% (ASI97) were selected. For each primer, the species coverage with no ASI97 (SC-NASI97) was calculated. Based on the SC-NASI97 parameter, the best primer pairs were OP_F053-KP_R020 for bacteria (region V1-V3; primer pair position for Escherichia coli J01859.1: 9-356); KP_F018-KP_R002 for archaea (V4; undefined-532); and OP_F114-KP_R031 for both (V3-V5; 340-801). Around 80% of the oral-bacteria and oral-archaea species analyzed had an ASI97 with at least one other species. These very similar species play different roles in the oral microbiota and belong to bacterial genera such as Campylobacter, Rothia, Streptococcus and Tannerella, and archaeal genera such as Halovivax, Methanosarcina and Methanosalsum. Moreover, ~20% and ~30% of these two-by-two similarity relationships were established between species from different bacterial and archaeal genera, respectively. Even taxa from distinct families, orders, and classes could be grouped in the same possible OTU. Consequently, regardless of the primer pair used, sequence clustering with a 97% similarity provides an inaccurate description of oral-bacterial and oral-archaeal species, which can greatly affect microbial diversity parameters. As a result, OTU clustering conditions the credibility of associations between some oral species and certain health and disease conditions. This significantly limits the comparability of the microbial diversity findings reported in oral microbiome literature.
Reliable and effective diagnostic systems are of vital importance for COVID-19, specifically for triage and screening procedures. In this work, a fully automatic diagnostic system based on chest X-ray images (CXR) has been proposed. It relies on the few-shot paradigm, which allows to work with small databases. Furthermore, three components have been added to improve the diagnosis performance: (1) a region proposal network which makes the system focus on the lungs; (2) a novel cost function which adds expert knowledge by giving specific penalties to each misdiagnosis; and (3) an ensembling procedure integrating multiple image comparisons to produce more reliable diagnoses. Moreover, the COVID-SC dataset has been introduced, comprising almost 1100 AnteroPosterior CXR images, namely 439 negative and 653 positive according to the RT-PCR test. Expert radiologists divided the negative images into three categories (normal lungs, COVID-related diseases, and other diseases) and the positive images into four severity levels. This entails the most complete COVID-19 dataset in terms of patient diversity. The proposed system has been compared with state-of-the-art methods in the COVIDGR-1.0 public database, achieving the highest accuracy (81.13% ± 2.76%) and the most robust results. An ablation study proved that each system component contributes to improve the overall performance. The procedure has also been validated on the COVID-SC dataset under different scenarios, with accuracies ranging from 70.81 to 87.40%. In conclusion, our proposal provides a good accuracy appropriate for the early detection of COVID-19.
Chronological age and biological sex estimation are two key tasks in a variety of procedures, including human identification and migration control. Issues such as these have led to the development of both semiautomatic and automatic prediction models, but the former are expensive in terms of time and human resources, while the latter lack the interpretability required to be applicable in real-life scenarios. This paper therefore proposes a new, fully automatic methodology for the estimation of age and sex. This first applies a tooth detection by means of a modified CNN with the objective of extracting the oriented bounding boxes of each tooth. Then, it feeds the image features inside the tooth boxes into a second CNN module designed to produce per-tooth age and sex probability distributions. The method then adopts an uncertainty-aware policy to aggregate these estimated distributions. Our approach yielded a lower mean absolute error than any other previously described, at 0.97 years. The accuracy of the sex classification was 91.82%, confirming the suitability of the teeth for this purpose. The proposed model also allows analyses of age and sex estimations on every tooth, enabling experts to identify the most relevant for each task or population cohort or to detect potential developmental problems. In conclusion, the performance of the method in both age and sex predictions is excellent and has a high degree of interpretability, making it suitable for use in a wide range of application scenarios.
Background: Zebrafish (Danio rerio) is a model organism for the study of human cancer. Compared with the murine model, the zebrafish model has several properties ideal for personalized therapies. The transparency of the zebrafish embryos and the development of the pigment-deficient "casper" zebrafish line give the capacity to directly observe cancer formation and progression in the living animal. Automatic quantification of cellular proliferation in vivo is critical to the development of personalized medicine. Methods: A new methodology was defined to automatically quantify the cancer cellular evolution. ZFTool was developed to establish a base threshold that eliminates the embryo autofluorescence, automatically measures the area and intensity of GFP (green-fluorescent protein) marked cells, and defines a proliferation index. Results: The proliferation index automatically computed on different targets demonstrates the efficiency of ZFTool to provide a good automatic quantification of cancer cell evolution and dissemination. Conclusion: Our results demonstrate that ZFTool is a reliable tool for the automatic quantification of the proliferation index as a measure of cancer mass evolution in zebrafish, eliminating the influence of its autofluorescence.
This in silico investigation aimed to: 1) evaluate a set of primer pairs with high coverage, including those most commonly used in the literature, to find the different oral species with 16S rRNA gene amplicon similarity/identity (ASI) values ≥97%; and 2) identify oral species that may be erroneously clustered in the same operational taxonomic unit (OTU) and ascertain whether they belong to distinct genera or other higher taxonomic ranks. Thirty-nine primer pairs were employed to obtain amplicon sequence variants (ASVs) from the complete genomes of 186 bacterial and 135 archaeal species. For each primer, ASVs without mismatches were aligned using BLASTN and their similarity values were obtained. Finally, we selected ASVs from different species with an ASI value ≥97% that were covered 100% by the query sequences. For each primer, the percentage of species-level coverage with no ASI≥97% (SC-NASI≥97%) was calculated. Based on the SC-NASI≥97% values, the best primer pairs were OP\_F053-KP\_R020 for bacteria (65.05%), KP\_F018-KP\_R002 for archaea (51.11%), and OP\_F114-KP\_R031 for bacteria and archaea together (52.02%). Eighty percent of the oral-bacteria and oralarchaea species shared an ASI≥97% with at least one other taxa, including Campylobacter , Rothia , Streptococcus , and Tannerella , which played conflicting roles in the oral microbiota. Moreover, around a quarter and a third of these two-by-two similarity relationships were between species from different bacteria and archaea genera, respectively. Furthermore, even taxa from distinct families, orders, and classes could be grouped in the same cluster. Consequently, irrespective of the primer pair used, OTUs constructed with a 97% similarity provide an inaccurate description of oral-bacterial and oral-archaeal species, greatly affecting microbial diversity parameters. As a result, clustering by OTUs impacts the credibility of the associations between some oral species and certain health and disease conditions. This limits significantly the comparability of the microbial diversity findings reported in oral microbiome literature. ### Competing Interest Statement The authors have declared no competing interest.