PURPOSE:This study addresses critical gaps in automated lymphoma segmentation from PET/CT imaging, often overlooked in prior work. While deep learning has been applied to this task, few studies evaluate generalizability on external or out-of-distribution data. Similarly, intra- and inter-observer variability analyses remain rare, limiting understanding of task difficulty. Moreover, most methods emphasize global segmentation metrics, neglecting lesion-level characteristics that are crucial for clinical decision-making. METHODS:We propose a clinically-relevant evaluation framework to assess four commonly used deep segmentation networks (ResUNet, SegResNet, DynUNet, SwinUNETR) on 611 PET/CT cases from multi-institutional datasets spanning varied lymphoma subtypes and lesion characteristics. In addition to the Dice similarity coefficient (DSC), we compute prediction errors on clinical lesion measures and analyze DSC performance as a function of these measures. Additionally, we use traditional lesion-specific detection criteria (1 and 2), providing insights into network's performance in identifying and localizing lesions respectively, and propose an additional Criterion 3 for segmenting lesions based on metabolic characteristics. Finally, we contextualize network performance by comparing it to expert human observers through intra- and inter-observer variability analyses. RESULTS:Networks perform best on large, metabolically active lesions. Their error patterns closely resemble those of expert annotators, while small and faint lesions remain challenging for both networks and physicians. CONCLUSION:Our clinically-relevant benchmarking framework enables more consistent and meaningful evaluation of lymphoma segmentation models, supporting robust decision-making in patient care. The approach is extensible to other architectures and disease types. Code is available at: https://github.com/microsoft/lymphoma-segmentation-dnn.
Background:Predicting whether an organ offer will be accepted for transplantation remains challenging for several reasons, including large offer volumes, highly imbalanced observations (more declines than acceptances), and lack of information about the human decision-making process. Offer acceptance models are used for risk-adjusted program evaluations and policy development, but there is a lack of literature on baselines and best practices for predictive applications. We compared a suite of machine learning models, feature sets, and sampling procedures to identify performance impacts when training offer acceptance prediction models. Methods:We evaluated several kidney offer acceptance models from logistic regression to gradient boosted trees that were trained on donor and candidate characteristics. We then selected the best-performing model and augmented training data with additional features (e.g., distance from the closest airport to the transplant hospital) or additional sampling procedures (e.g., undersampling). Results:Compared to the baseline logistic regression model (average precision: 0.0645), the XGBoost model offered the best performance improvement over the baseline (average precision: 0.0907). Including transportation-related features in the model further improved model performance (average precision: 0.0940); however, we did not observe substantial model performance differences based on the sampling procedure used. Conclusions:Leveraging advanced machine learning models and incorporating nonclinical datapoints (like transportation distances) can improve transplant organ offer acceptance prediction models. However, we observed steep tradeoffs between precision and recall as captured in the low average precision scores despite deceptively high AUROCs (baseline AUROC 0.832). Our findings suggest that even the best-performing models would not provide clear, equitable benefits over existing allocation policies. More research is needed before these models are practical for clinical implementation.
Ear disease contributes significantly to global hearing loss, with recurrent otitis media being a primary preventable cause in children, impacting development. Artificial intelligence (AI) offers promise for early diagnosis via otoscopic image analysis, but dataset biases and inconsistencies limit model generalizability and reliability. This retrospective study systematically evaluated three public otoscopic image datasets (Chile; Ohio, USA; Türkiye) using quantitative and qualitative methods. Two counterfactual experiments were performed: (1) obscuring clinically relevant features to assess model reliance on non-clinical artifacts, and (2) evaluating the impact of hue, saturation, and value on diagnostic outcomes. Quantitative analysis revealed significant biases in the Chile and Ohio, USA datasets. Counterfactual Experiment I found high internal performance (AUC > 0.90) but poor external generalization, because of dataset-specific artifacts. The Türkiye dataset had fewer biases, with AUC decreasing from 0.86 to 0.65 as masking increased, suggesting higher reliance on clinically meaningful features. Counterfactual Experiment II identified common artifacts in the Chile and Ohio, USA datasets. A logistic regression model trained on clinically irrelevant features from the Chile dataset achieved high internal (AUC = 0.89) and external (Ohio, USA: AUC = 0.87) performance. Qualitative analysis identified redundancy in all the datasets and stylistic biases in the Ohio, USA dataset that correlated with clinical outcomes. In summary, dataset biases significantly compromise reliability and generalizability of AI-based otoscopic diagnostic models. Addressing these biases through standardized imaging protocols, diverse dataset inclusion, and improved labeling methods is crucial for developing robust AI solutions, improving high-quality healthcare access, and enhancing diagnostic accuracy.
Here, we summarize the work that Microsoft's philanthropic Artificial Intelligence (AI) for Good Lab has completed in the realm of promoting public and population health. In particular, after providing examples of how the AI for Good Lab has articulated the value of using AI to improve public and population health, we provide examples and references of the work demonstrating how the Lab has: applied Artificial Intelligence (AI) to improve maternal, fetal, and infant health; leveraged large language models to improve population health; and applied AI to improve rural health and healthcare. We also summarize what we have learned through our work, finding that: getting the question right and ensuring the limitations of any analysis are understood is important; collaboration across public, private, and educational institutions with subject matter experts will be the most effective and efficient way to harness this new technology; and that focusing on metrics that reflect health, and not just the accuracy of the model, is the most impactful way to improve the health of populations, worldwide.
Microsoft Research, United States. Address correspondence and reprint requests to James N. Weinstein, DO, MS, Microsoft Research, United States; E-mail: [email protected] Received 9 September, 2020 Accepted 11 September, 2020 The manuscript submitted does not contain information about medical device(s)/drug(s). No funds were received in support of this work. No relevant financial activities outside the submitted work.
Importance:Retinopathy of prematurity (ROP) is the leading cause of preventable childhood blindness worldwide. If detected and treated early, ROP-associated blindness is preventable; however, identifying patients who might respond to treatment requires screening over time, which is challenging in low-resource settings where access to pediatric ophthalmologists and pediatric ocular imaging cameras is limited. Objective:To develop and assess the performance of a machine learning algorithm that uses smartphone-collected videos to perform retinal screening for ROP in low-resource settings. Design, Setting, and Participants:This diagnostic study used smartphone-obtained videos of fundi in premature neonates with and without ROP in Mexico and Argentina between May 12, 2020, and October 31, 2023. Machine-learning (ML)-driven algorithms were developed to process a video, identify the best frames within the video, and use those frames to determine whether ROP was likely or not. Eligible neonates born with gestational age less than 36 weeks or birth weight less than 1500 g were included on the study. Exposures:An ML algorithm applied to a smartphone-obtained video. Main Outcomes and Measures:The ML algorithms' ability to identify high-quality retinal images and classify those images as indicating ROP or not at the frame and patient levels, measured by accuracy, specificity, and sensitivity, compared with classifications from 3 pediatric ophthalmologists. Results:A total of 524 videos were collected for 512 neonates with median gestational age of 32 weeks (range, 25-36 weeks) and median birth weight of 1610 g (range, 580-2800 g). The frame selection model identified high-quality retinal images from 397 of 456 videos (87.1%; 95% CI, 84.0%-90.1%) reserved for testing model performance. Across all test videos, 97.4% (95% CI, 96.7%-98.1%) of high-quality retinal images selected by the model contained fundus images. At the frame level, the ROP classifier model had a sensitivity of 76.7% (95% CI, 69.9%-83.5%); at the patient level, the classifier model had a sensitivity of 93.3% (95% CI, 86.4%-100%). At both levels, the model's sensitivity was higher than that for the panel of pediatric ophthalmologists (frame level: 71.4% [95% CI, 64.1%-78.7%]; patient level: 73.3% [95% CI, 61.0%-85.6%]). Specificity and accuracy were higher for ophthalmologist classification vs the ML model. Conclusions and Relevance:In this diagnostic study, a process that used smartphone-collected videos of premature neonates' fundi to determine whether high-quality retinal images were present had high sensitivity to classify such images as indicating or not indicating ROP but lower specificity and accuracy than ophthalmologist assessment. This process costs a fraction of the current process for retinal image collection and classification and could be used to expand access to ROP screening in low-resource settings, with potential to help prevent the most common cause of preventable childhood blindness.
Purpose: To assess the performance of general-domain large language models (LLMs), particularly OpenAI’s Generative Pre-trained Transformer (GPT) models, within the American Academy of Ophthalmology (AAO) Self-Assessment Program, which is based on AAO’s Basic and Clinical Science Course. Methods: We input 3357 questions into GPT-4o, GPT-4-Turbo, o1 and o3-mini via Microsoft’s Azure OpenAI Service using zero-shot and chain-of-thought (CoT) prompting. Questions with images were analyzed using the multimodal version of GPT-4o and GPT-4.1. The performance of the LLMs was compared to 1371 unique residents who had previously participated in the program. Additionally, we compared the performance on 1399 questions, including information on 3 question types: recall, interpretation, and decision-making or clinical management. Average accuracy rates were used to evaluate performance and compare statistical significance across categories. Results: o1 (CoT) was the most accurate model (95% confidence interval [CI]: 90.3%–92.1%) with performance ranging from 95.17% (general medicine) to 86.9% (cornea) and 91.1% accuracy on a synthesized sample test. It also outperformed residents in recall-type, interpretation-type, and decision-making or clinical management questions (95.7%, 85.3%, and 90.8%, respectively, P < 0.001). Third-year residents were more accurate than first-year or second-year residents (78.2%, 68.3%, 74.9%, respectively). On multimodal inputs, adding images improved the model’s accuracy but all models still underperformed compared to residents. Conclusions: The accuracy of the LLMs models continues to improve, with o1 (CoT) showing the highest overall performance. Multimodal inputs can enhance model accuracy, but current models still need improvement. LLMs shows great potential in democratizing access to high-quality medical knowledge.
Objective: To investigate an ensemble-based approach utilizing deep learning models for accurate and interpretable detection of macular telangiectasia (MacTel) type 2 on OCT imaging. Design: Retrospective analysis of OCT scans, model development, and assessment. Participants: A total of 5200 OCT images from participants in the MacTel Registry conducted by the Lowy Medical Research Institute and from the University of Washington (780 MacTel patients and 1900 non-MacTel patients). Methods, Intervention, or Testing: We trained multiple individual MacTel vs. non-MacTel classification models using traditional supervised learning and self-supervised learning (SSL) and ensembled them using average weighting methods. We investigated diverse methodologies for constructing the ensemble, including varied architectural configurations and learning paradigms of individual models, and manipulating the amount of labeled data accessible for training. Model performance was compared against human expert graders on held-out test set data. Model interpretability was investigated using gradient-weighted class activation maps (Grad-CAM) visualization and by evaluating interrater agreement. Main Outcome Measures: For model performance, area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), accuracy, sensitivity, and specificity were reported. For interpretability, interrater agreements and Grad-CAM visualization results were evaluated. Results: Despite access to only 419 OCT volumes, including 185 MacTel patients within the 10% labeled training dataset, the ensemble model demonstrated a performance level (AUROC 0.972 [95% confidence interval (CI), 0.971-0.973], AUPRC 0.967 [95% CI, 0.965-0.969], accuracy 91.7%, sensitivity 0.905, and specificity 0.925) comparable to the human experts ensemble (AUROC 0.977 [95% CI, 0.975-0.978], AUPRC 0.987 [95% CI, 0.986-0.987], accuracy 96.8%, sensitivity 0.929, and specificity 1) on a test set of 500 patients. The individual models did not achieve the same performance levels when evaluated separately. Conclusions: Even with limited data, combining SSL with ensemble approaches improved MacTel classification accuracy and interpretation compared to the individual models. Self-supervised learning captures meaningful representations from unlabeled data, a key benefit in the setting of limited data such as with rare diseases.
Image-to-image (I2I) translation networks have emerged as promising tools for generating synthetic medical images; however, their clinical reliability and ability to preserve diagnostically relevant features remain underexplored. This study evaluates the performance of state-of-the-art 2D/3D I2I networks for converting ultrasound (US) images to synthetic MRI in prostate cancer (PCa) imaging. The novelty lies in combining radiomics, expert clinical evaluation, and classification performance to comprehensively benchmark these models for potential integration into real-world diagnostic workflows. A dataset of 794 PCa patients was analyzed using ten leading I2I networks to synthesize MRI from US input. Radiomics feature (RF) analysis was performed using Spearman correlation to assess whether high-performing networks (SSIM > 0.85) preserved quantitative imaging biomarkers. A qualitative evaluation by seven experienced physicians assessed the anatomical realism, presence of artifacts, and diagnostic interpretability of synthetic images. Additionally, classification tasks using synthetic images were conducted using two machine learning and one deep learning model to assess the practical diagnostic benefit. Among all networks, 2D-Pix2Pix achieved the highest SSIM (0.855 ± 0.032). RF analysis showed that 76 out of 186 features were preserved post-translation, while the remainder were degraded or lost. Qualitative feedback revealed consistent issues with low-level feature preservation and artifact generation, particularly in lesion-rich regions. These evaluations were conducted to assess whether synthetic MRI retained clinically relevant patterns, supported expert interpretation, and improved diagnostic accuracy. Importantly, classification performance using synthetic MRI significantly exceeded that of US-based input, achieving average accuracy and AUC of 0.93 ± 0.05. Although 2D-Pix2Pix showed the best overall performance in similarity and partial RF preservation, improvements are still required in lesion-level fidelity and artifact suppression. The combination of radiomics, qualitative, and classification analyses offered a holistic view of the current strengths and limitations of I2I models, supporting their potential in clinical applications pending further refinement and validation.
Artificial intelligence (AI) can advance medical diagnostics, but interpretability limits its clinical use. This work links standardized quantitative Radiomics features (RF) extracted from medical images with clinical frameworks like PI-RADS, ensuring AI models are understandable and aligned with clinical practice. We investigate the connection between visual semantic features defined in PI-RADS and associated risk factors, moving beyond abnormal imaging findings, and establishing a shared framework between medical and AI professionals by creating a standardized radiological/biological RF dictionary. Six interpretable and seven complex classifiers, combined with nine interpretable feature selection algorithms (FSA), were applied to RFs extracted from segmented lesions in T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI), and apparent diffusion coefficient (ADC) multiparametric MRI sequences to predict TCIA-UCLA scores, grouped as low-risk (scores 1-3) and high-risk (scores 4-5). We then utilized the created dictionary to interpret the best predictive models. Combining sequences with FSAs including ANOVA F-test, Correlation Coefficient, and Fisher Score, and utilizing logistic regression, identified key features: The 90th percentile from T2WI, (reflecting hypo-intensity related to prostate cancer risk; Variance from T2WI (lesion heterogeneity; shape metrics including Least Axis Length and Surface Area to Volume ratio from ADC, describing lesion shape and compactness; and Run Entropy from ADC (texture consistency). This approach achieved the highest average accuracy of 0.78 +/- 0.01, significantly outperforming single-sequence methods (p-value < 0.05). The developed dictionary for Prostate-MRI (PM1.0) serves as a common language and fosters collaboration between clinical professionals and AI developers to advance trustworthy AI solutions that support reliable/interpretable clinical decisions.
Background Female physicians have lower incomes than male physicians. While overall sex-based income disparities are dramatic, compensation differs considerably across specialties. A better understanding of the relationship between anticipated specialty-specific annual incomes and the proportion of females entering that specialty might help residency program directors argue for equity in specialty choice and income for female physicians. Objective We sought to determine the relationship between the percentage of females in the workforce entering a specialty and the average compensation of that specialty in 2023. Methods From a recent JAMA article, we obtained the characteristics and numbers of trainees engaged in 160 specialties or subspecialties within 13 489 graduate medical education programs in 2023; we aggregated those data into 50 specialties for which 2023 average annual self-reported compensation were publicly available from Doximity. We conducted a stepwise linear regression in which the specialty-specific proportion of trainees who were female, US medical school graduates, Canadian, Doctors of Osteopathy, American Indian or Alaska Native, Asian, Black, Hispanic or Latino, Pacific Islander, or White were used to predict the specialty-specific average annual income. We conducted the analysis in 2024. Results Each one percent increase in the specialty-specific percentage of female trainees was associated with a $5,301 decrease in average specialty-specific annual compensation, and each one percent increase in the US medical school graduate percentage was associated with a $3,821 increase. These 2 characteristics accounted for 78% of the adjusted explainable variance in average specialty-specific annual compensation. Conclusions Specialties with higher proportions of female trainees had lower average annual compensation rates.
Importance:Multiple choice questions (MCQs) are an important and integral component of ophthalmology residency training evaluation and board certification; however, high-quality questions are difficult and time-consuming to draft. Objective:To evaluate whether general-domain large language models (LLMs), particularly OpenAI's Generative Pre-trained Transformer 4 (GPT-4), can reliably generate high-quality, novel, and readable MCQs comparable to those of a committee of experienced examination writers. Design, Setting, and Participants:This survey study, conducted from September 2024 to April 2025, assesses LLM performance in generating MCQs based on the American Academy of Ophthalmology (AAO) Basic and Clinical Science Course (BCSC) compared with a committee of human experts. Ten expert ophthalmologists, who were masked to the generation source, independently evaluated MCQs using a 10-point Likert scale (1 = extremely poor; 10 = criterion standard quality) across 5 criteria: appropriateness, clarity and specificity, relevance, discriminative power, and suitability for trainees. Intervention:Relevant BCSC content and AAO question-writing guidelines were input into GPT-4o via Microsoft's Azure OpenAI Service, and structured prompts were used to generate MCQs. Main Outcomes and Measures:The primary outcomes were median scores and statistical comparisons using the bootstrapping method; string similarity scores based on Levenshtein distance (0-100, with 100 indicating identical content) between LLM-MCQs and the entire BCSC question bank; Flesch Reading Ease metric for readability; and intraclass correlation coefficient (ICC) for inter-rater agreement are reported. Results:The 10 graders had between 1 and 28 years of clinical experience in ophthalmology (median [IQR] experience, 6 years [3-15 years]). Questions generated by GPT-4 and a committee of experts received median scores of 9 and 9 in combined scores, appropriateness, clarity and specificity, and relevance (difference, 0; 95% CI, 0-0; P > .99); 8 and 9 in discriminative power (difference, 1; 95% CI, -1 to 1; P = .52); and 8 and 8 in suitability for trainees (difference, 0; 95% CI, -1 to 0; P > .99), respectively. Nearly 95% of LLM-MCQs had similarity scores less than 60, indicating most LLM-MCQs had limited or no resemblance to existing content. Interrater reliability was moderate (ICC, 0.63; P < .001), and mean (SD) readability scores were similar across sources (37.14 [22.54] vs 42.60 [22.84]; P > .99). Conclusions and Relevance:In this survey study, results indicate that an LLM could be used to develop ophthalmology board-style MCQs and expand examination banks to further support ophthalmology residency training. Despite most questions having a low similarity score, the quality, novelty, and readability of the LLM-generated questions need to be further assessed.
Importance:Deep learning image analysis often depends on large, labeled datasets, which are difficult to obtain for rare diseases. Objective:To develop a self-supervised approach for automated classification of macular telangiectasia type 2 (MacTel) on optical coherence tomography (OCT) with limited labeled data. Design, Setting, and Participants:This was a retrospective comparative study. OCT images from May 2014 to May 2019 were collected by the Lowy Medical Research Institute, La Jolla, California, and the University of Washington, Seattle, from January 2016 to October 2022. Clinical diagnoses of patients with and without MacTel were confirmed by retina specialists. Data were analyzed from January to September 2023. Exposures:Two convolutional neural networks were pretrained using the Bootstrap Your Own Latent algorithm on unlabeled training data and fine-tuned with labeled training data to predict MacTel (self-supervised method). ResNet18 and ResNet50 models were also trained using all labeled data (supervised method). Main Outcomes and Measures:The ground truth yes vs no MacTel diagnosis is determined by retinal specialists based on spectral-domain OCT. The models' predictions were compared against human graders using accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), area under precision recall curve (AUPRC), and area under the receiver operating characteristic curve (AUROC). Uniform manifold approximation and projection was performed for dimension reduction and GradCAM visualizations for supervised and self-supervised methods. Results:A total of 2636 OCT scans from 780 patients with MacTel and 131 patients without MacTel were included from the MacTel Project (mean [SD] age, 60.8 [11.7] years; 63.8% female), and another 2564 from 1769 patients without MacTel from the University of Washington (mean [SD] age, 61.2 [18.1] years; 53.4% female). The self-supervised approach fine-tuned on 100% of the labeled training data with ResNet50 as the feature extractor performed the best, achieving an AUPRC of 0.971 (95% CI, 0.969-0.972), an AUROC of 0.970 (95% CI, 0.970-0.973), accuracy of 0.898%, sensitivity of 0.898, specificity of 0.949, PPV of 0.935, and NPV of 0.919. With only 419 OCT volumes (185 MacTel patients in 10% of labeled training dataset), the ResNet18 self-supervised model achieved comparable performance, with an AUPRC of 0.958 (95% CI, 0.957-0.960), an AUROC of 0.966 (95% CI, 0.964-0.967), and accuracy, sensitivity, specificity, PPV, and NPV of 90.2%, 0.884, 0.916, 0.896, and 0.906, respectively. The self-supervised models showed better agreement with the more experienced human expert graders. Conclusions and Relevance:The findings suggest that self-supervised learning may improve the accuracy of automated MacTel vs non-MacTel binary classification on OCT with limited labeled training data, and these approaches may be applicable to other rare diseases, although further research is warranted.
Introduction Diabetes is a leading contributor to cardiovascular disease and mortality; social determinants of health (SDOH) are associated with disparities in diabetes risk. Quantifying the cumulative impact of SDOH and identifying the SDOH most associated with diabetes prevalence at the neighbourhood level can help policy-makers design and target local interventions to mitigate these disparities. Machine learning (ML) methods can provide novel insights and help inform public health intervention strategies in a place-based manner.Methods In a cross-sectional study, we used gradient boosting ML models to estimate the cumulative contribution of a set of SDOH variables to diabetes prevalence (%) at the census tract level within New York City (NYC); Shapley Additive Explanations were used to assess the magnitude and shape of relationships between our SDOH variables and model-predicted NYC diabetes prevalence. SDOH measures included socioeconomic position, educational attainment, food access, air quality, neighbourhood environment, housing conditions and insurance coverage.Results Across 2096 NYC census tracts (population 8 170 505), mean diabetes prevalence was 11.5% (SD 3.7%; range 1.9%–42.8%). A set of 16 SDOH variables representing a framework of 16 distinct SDOH concepts accounted for 67% of the between-tract variance in model-derived NYC diabetes prevalence estimates (95% CI 66% to 68%); a set of 81 variables representing these 16 concepts accounted for 80% of variance (95% CI 78% to 81%). Models showed excellent across-location generalisation. The most important variables driving model predictions within NYC were measures of low educational attainment and poverty.Conclusions SDOH accounted for a substantial proportion of neighbourhood-level variation in diabetes prevalence within NYC, independent of the demographics and health behaviours associated with those SDOH. Our place-based findings suggest that, within NYC, where approximately one million residents have diabetes and there are legislative requirements to reduce the impacts from diabetes, policies reducing socioeconomic and educational inequality could have the greatest potential to equitably achieve this.
My appreciation for art led me to Gustav Klimt, who, in 1901, presented a controversial painting titled "Medicine." The goddess Hygeia, who symbolized health, hygiene, and well-being, appears in the center; she stands between a beautiful woman representing the river of life and an abundance of people in various states of despair representing the suffering often associated with illness. Klimt presented this painting to the medicine faculty at the University of Vienna, but they found it offensive, as it implicitly acknowledged their inadequacies in addressing the health needs of the disenfranchised.
This study investigates the foundational characteristics of image-to-image translation networks, specifically examining their suitability and transferability within the context of routine clinical environments, despite achieving high levels of performance, as indicated by a Structural Similarity Index (SSIM) exceeding 0.95. The evaluation study was conducted using data from 794 patients diagnosed with Prostate cancer. To synthesize MRI from Ultrasound images, we employed five widely recognized image to image translation networks in medical imaging: 2DPix2Pix, 2DCycleGAN, 3DCycleGAN, 3DUNET, and 3DAutoEncoder. For quantitative assessment, we report four prevalent evaluation metrics Mean Absolute Error, Mean Square Error, Structural Similarity Index (SSIM), and Peak Signal to Noise Ratio. Moreover, a complementary analysis employing Radiomic features (RF) via Spearman correlation coefficient was conducted to investigate, for the first time, whether networks achieving high performance, SSIM greater than 0.9, could identify low-level RFs. The RF analysis showed 76 features out of 186 RFs were discovered via just 2DPix2Pix algorithm while half of RFs were lost in the translation process. Finally, a detailed qualitative assessment by five medical doctors indicated a lack of low level feature discovery in image to image translation tasks.
OPINION article Front. Artif. Intell., 19 June 2024Sec. Medicine and Public Health Volume 7 - 2024 | https://doi.org/10.3389/frai.2024.1430756