Dermatologic imaging has been rapidly expanding, with over 70% of related PubMed articles published since 2016 and over a million images across international research challenges and large-scale datasets with skin images. To improve data quality and usability, standardizing dermatologic imaging data for non-protected health information (non-PHI) research systems is essential. While the International Skin Imaging Collaboration (ISIC) has advanced standards in skin imaging, the field lacks a generalizable infrastructure to organize and describe imaging data for non-PHI research systems. This results in inconsistently labeled, heterogeneous datasets that hinder data integration, scalability, and interoperability. To address this gap, we propose the Dermatology Imaging Data Structure (DermIDS), inspired by the Brain Imaging Data Structure (BIDS) for neuroimaging. This structured framework aims to improve usability across datasets, reveal metadata gaps, and enable scalable artificial intelligence (AI)/machine learning (ML)-ready workflows. To illustrate this system, we curated and processed 1,000,692 images with DermIDS. We demonstrate that DermIDS (1) supports multimodal photographic data acquired from clinical photography, general photography, dermoscopy, reflectance confocal microscopy, and surface 3D imaging; (2) facilitates image-specific technical and clinical metadata organization; and (3) streamlines quality control and harmonization. Across all images, 1,256 unique metadata features were identified. However, 70% of clinical metadata features and 98% of technical metadata features were present in less than 100k images, highlighting key gaps and demonstrating the utility of DermIDS in revealing inconsistencies and opportunities for standardization. Our work supports large-scale analysis and harmonization, laying the foundation for AI/ML-ready workflows to advance dermatologic imaging research.
Skin segmentation from clinical photography is a crucial step in dermatological image analysis. However, the variability in skin tones, lighting conditions, anatomical regions, and the presence of additional objects introduces significant challenges. Due to these complexities, the segmentation process is often performed manually, as developing an algorithm capable of handling such diverse conditions is particularly difficult. Recently, openworld foundation models have emerged, offering the potential to generalize across diverse and unseen conditions. These models present a promising opportunity for dermatology. In this work, we adopt two such models-Grounding DINO and SAM 2-to construct a pipeline for zero-shot skin segmentation in dermatology. We evaluated our approach on two clinical skin photography datasets comprising 27,378 images. Based on a manual rating protocol, 77.1% of the segmentations were deemed acceptable, demonstrating robustness in handling realworld clinical photographs. Our results highlight the potential of open-world foundation models to address a challenging problem in dermatology with minimal human involvement.
Abstract: There is an urgent need for validated tools to measure sclerotic cutaneous chronic graft-versus-host disease (scGVHD). We examined the interobserver reproducibility within a session and intraobserver repeatability between sessions of the Myoton device for quantifying skin sclerosis in 36 adults with scGVHD. The Myoton was used to measure oscillation frequency and relaxation time of soft tissues at 7 bilateral sites (14 anatomic sites) by 2 study personnel at 2 study sessions. Agreement was measured using mean pairwise absolute difference (MPD), and reliability was measured using intraclass correlation coefficient (ICC). For each of the 2 Myoton parameters, the overall interobserver MPD was <5% of the average overall values and the interobserver ICC was >0.90 between the 2 observers, indicating excellent agreement and reliability within a measurement session. The median time between sessions 1 and 2 was 47.5 days. The overall normalized intraobserver MPD was <7% of the average overall values for each of the 2 Myoton parameters, reflecting good agreement between sessions. The intraobserver ICC for frequency and relaxation time parameters were 0.85 and 0.84, respectively, indicating good reliability between sessions. The reproducibility and repeatability of a bonus site selected at each study visit were similar to the standard 14 anatomic sites. However, no individual site was nearly as reproducible or repeatable as the overall Myoton measurements averaged across the patient. Our findings emphasize the utility of the Myoton for assessing skin properties in scGVHD with patient-level measurements.
We noninvasively visualized the upper dermal microvasculature of 31 patients after hematopoietic cell transplantation (HCT) by reflectance confocal videomicroscopy. The microvessel diameter and number and diameter of adherent and rolling leukocytes for patients after HCT were similar to historically published values in healthy subjects. We also observed "paused" leukocytes i.e. leukocytes that temporarily stop, coinciding with the simultaneous stopping of the rest of the blood flow. The number and diameter of paused leukocytes, and the duration of leukocyte being paused for patients after HCT were also similar to historically published values in healthy subjects. However, we observed more blood vessels per imaging field of view (500 × 500 μm2) in the skin of patients after HCT than healthy subjects (a median of 3 versus 2). The number of blood vessels in a field of view was not correlated with the number of adherent and rolling leukocytes. Vessel size (flow width) had a meaningful correlation with the diameter, but not number, of paused leukocytes. Paused leukocyte diameter had no correlation with the duration of pausing. Reflectance confocal videomicroscopy enables characterization of intact upper dermal microvasculature of patients with extremely altered immune system.
We designed and successfully implemented photography protocols for the PALM007 (Pamoja Tulinde Maisha - Together Save Lives in Swahili) randomised clinical trial evaluating tecovirimat for the treatment of mpox in the Democratic Republic of the Congo (DRC). We addressed the unique challenges and limitations of conducting standardised photography procedures on patients with mpox in high-risk 'red zones' in remote study treatment centre sites. Key considerations included standardised photography protocols, training, quality assurance, participant privacy and comfort, and interpersonal communication. Our developed procedures enabled acquisition of 61 926 standardised, clinical-quality photographs of 597 patients with mpox. These considerations may guide future studies collecting standardised photographs in a low-resource setting.
The National Institutes of Health (NIH) chronic graft-versus-host disease (cGVHD) Consensus Project established response criteria that enabled clinical trials and facilitated regulatory approval of multiple therapies. Nonetheless, organ-specific assessments have limitations, particularly for severe sclerotic skin involvement. Cutaneous manifestations occur in approximately half of patients with chronic GVHD but are difficult to evaluate due to heterogeneous clinical features. To address these challenges, the NIH Consensus Skin Task Force convened from 2024 to 2025 to refine skin response measures for use in clinical trials. This report (1) summarizes current diagnosis and scoring of skin chronic GVHD, (2) reviews existing response assessments, (3) identifies gaps in their performance, (4) proposes refinements to the 2014 NIH skin response criteria, and (5) outlines future directions incorporating patient-reported outcomes, novel technologies, and biomarker research. Current NIH scoring relies on a 4-point body surface area (BSA) scale and a 3-point sclerosis features scale. While standardized, these measures have limited sensitivity to clinically meaningful change. Proposed refinements include: (1) separate BSA assessments for epidermal involvement and sclerotic features; (2) modification of sclerosis descriptors, with removal of impaired mobility and specification of GVHD-related ulceration, and addition of sclerosis-associated edema with or without erythema; and (3) replacement of the exploratory severity scale with two new clinician instruments. The Sclerosis Quality and Physical Signs scale (0 to 10) captures qualitative physical changes, while the Sclerosis Daily Function Impact scale (0 to 4) evaluates functional compromise. These proposed refinements aim to improve the accuracy, reproducibility, and clinical relevance of skin chronic GVHD assessments, strengthening trial endpoints and patient care.
Standardized 2D photography plays an essential role in dermatologic practice, supporting longitudinal documentation, patient monitoring, and consensus-based clinical scoring. However, photographs taken from limited views often suffer from reduced anatomical context and missing body-location information. 3D representations enable unified spatial interpretation of multi-view imagery. Recent developments in computer vision have made it feasible to infer dense correspondences between 2D images and a 3D human mesh. In this study, we explored integrating 2D dermatological images with a 3D surface model using DensePose, a deep learning-based human dense correspondence framework. This creates an anatomically grounded representation that supports mesh-level analyses and recovers spatial context for each image. We use a dataset including four full body photographs (front, back, and each side) from each of 147 subjects with chronic graft-versus-host disease, for a total of 588 images. Our method integrates these multiple 2D full-body photographs captured across varied body shapes and camera angles into a 3D mesh. We further showed that the resulting 3D mesh enables quantification of the extent to which individual 2D images, or their combinations, represent the complete body surface. On average, a single full body view captures 28% of the body surface, while adding a second, third, and fourth view increases average coverage to 50%, 72%, and 80%, respectively. To assess spatial consistency, we annotated up to 10 anatomical landmarks per patient on 80 images across 20 patients and reported a median pairwise geodesic distance between corresponding landmarks of 4.6 cm. These findings can guide how dermatology images are captured and support future opportunities in monitoring, education, and communication using existing infrastructure.
Mpox is a viral illness with heavy cutaneous involvement. Automatic tracking of mpox lesion progression is critical in determining the resolution of evolving lesions. This work introduces a novel application of deep learning for lesion monitoring through alignment of dermatological hand photographs. By adapting the VoxelMorph framework for 2D photographic data, we explore key point alignment across serial images. We trained our neural network model on a unique dataset of 1,658 hand images and evaluated its performance on a test set of 254 images. Additionally, we validated the method's generalizability with a supplementary set of 500 images, which included extensive Mpox infection. Our findings indicate modest yet significant improvements in key points and lesion center registration across different regularization strengths. Although promising, the complexity of hand structure presents challenges, requiring cautious application and further refinement, especially in regions with intense spatial discontinuities, such as interdigital areas.
Measuring skin involvement in chronic graft-versus-host disease (cGVHD) currently requires expert manual assessment, which is costly, time-consuming, and shows high interrater disagreement (>20% surface area). In our previous work, automated image analysis showed promise for measuring affected skin area under controlled photography conditions. Our aim is to improve the performance of these methods in standard clinical photographs without the need for costly expert annotations using a semi-supervised approach. A baseline U-Net model was trained in a fully supervised manner using 360 3D photographs from 36 cGVHD patients, with expert-marked ground truth contours of affected skin. The model was then iteratively retrained by incorporating an additional 5648 unlabeled photographs from 83 new patients using a semi-supervised method. Testing on clinical photographs of 20 held-out patients, the median surface area error improved from 19.2% (interquartile range 6.3 - 33.8) at baseline to 10.2% (4.5 - 22.6) after retraining. Semi-supervised training therefore provides an effective method for translating a pre-trained U-Net segmentation model to standard clinical photographs, without the need for additional expert annotations. Such models could help standardize cGVHD assessment and tracking, alleviating the need for costly expert evaluations and providing a reliable tool that would significantly enhance the current standard of manual assessment.
Optical coherence tomography (OCT) has sufficient depth penetration for detection of skin pathologies, but its detection effectiveness can be aided by the assistance of artificial intelligence (AI) modeling. AI model-building identifies pathologies by comparing images from healthy and diseased tissues, but healthy skin can present as quite variable across skin types and ages. Here, we selected a commonly used parameter for skin analysis and attenuation coefficient and analyzed how it varied in the dermis and epidermis, and in skin-exposed and skin-protected regions, for 100 subjects from a wide range of skin types (Fitzpatrick types I–V) and ages (13–83). For the statistical analysis, we report whether comparisons of the dermis and epidermis and sun-exposed and sun-protected areas across age and skin type are statistically significant, indeterminate, or not statistically significant and present 95% confidence intervals for this parameter as it ranges across different ages and skin types. This process of pre-analyzing features using healthy images provides a roadmap for how to ease the recruitment process while acquiring a sufficient range of images for effective AI model-building. We expect this type of analysis can have the effect of accelerating translation of AI-based OCT image analysis to the clinic.
Purpose Mpox is a viral illness with symptoms similar to smallpox. A key clinical metric to monitor disease progression is the number of skin lesions. Manually counting mpox skin lesions is labor-intensive and susceptible to human error. Approach We previously developed an mpox lesion counting method based on the UNet segmentation model using 66 photographs from 18 patients. We have compared four additional methods: the instance segmentation methods Mask R-CNN, YOLOv8, and E2EC, in addition to a UNet++ model. We designed a patient-level leave-one-out experiment, assessing their performance using F1 score and lesion count metrics. Finally, we tested whether an ensemble of the networks outperformed any single model. Results Mask R-CNN model achieved an F1 score of 0.75, YOLOv8 a score of 0.75, E2EC a score of 0.70, UNet++ a score of 0.81, and baseline UNet a score of 0.79. Bland-Altman analysis of lesion count performance showed a limit of agreement (LoA) width of 62.2 for Mask R-CNN, 91.3 for YOLOv8, 94.2 for E2EC, and 62.1 for UNet++, with the baseline UNet model achieving 69.1. The ensemble showed an F1 score performance of 0.78 and LoA width of 67.4. Conclusions Instance segmentation methods and UNet-based semantic segmentation methods performed equally well in lesion counting. Furthermore, the ensemble of the trained models showed no performance increase over the best-performing model UNet, likely because errors are frequently shared across models. Performance is likely limited by the availability of high-quality photographs for this complex problem, rather than the methodologies used.
Objectives/Goals: Manual skin assessment in chronic graft-versus-host disease (cGVHD) can be time consuming and inconsistent (>20% affected area) even for experts. Building on previous work we explore methods to use unmarked photos to train artificial intelligence (AI) models, aiming to improve performance by expanding and diversifying the training data without additional burden on experts. Methods/Study Population: Common to many medical imaging projects, we have a small number of expert-marked patient photos (N = 36, n = 360), and many unmarked photos (N = 337, n = 25,842). Dark skin (Fitzpatrick type 4+) is underrepresented in both sets; 11% of patients in the marked set and 9% in the unmarked set. In addition, a set of 20 expert-marked photos from 20 patients were withheld from training to assess model performance, with 20% dark skin type. Our gold standard markings were manual contours around affected skin by a trained expert. Three AI training methods were tested. Our established baseline uses only the small number of marked photos (supervised method). The semi-supervised method uses a mix of marked and unmarked photos with human feedback. The self-supervised method uses only unmarked photos without any human feedback. Results/Anticipated Results: We evaluated performance by comparing predicted skin areas with expert markings. The error was given by the absolute difference between the percentage areas marked by the AI model and expert, where lower is better. Across all test patients, the median error was 19% (interquartile range 6 – 34) for the supervised method and 10% (5 – 23) for the semi-supervised method, which incorporated unmarked photos from 83 patients. On dark skin types, the median error was 36% (18 – 62) for supervised and 28% (14 – 52) for semi-supervised, compared to a median error on light skin of 18% (5 – 26) for supervised and 7% (4 – 17) for semi-supervised. Self-supervised, using all 337 unmarked patients, is expected to further improve performance and consistency due to increased data diversity. Full results will be presented at the meeting. Discussion/Significance of Impact: By automating skin assessment for cGVHD, AI could improve accuracy and consistency compared to manual methods. If translated to clinical use, this would ease clinical burden and scale to large patient cohorts. Future work will focus on ensuring equitable performance across all skin types, providing fair and accurate assessments for every patient.
Optical coherence tomography (OCT) has sufficient depth penetration for detection of skin pathologies, but its detection effectiveness can be aided by the assistance of artificial intelligence (AI) modeling. AI model-building identifies pathologies by comparing images from healthy and diseased tissues, but healthy skin can present as quite variable across skin types and ages. Here, we selected a commonly used parameter for skin analysis and attenuation coefficient and analyzed how it varied in the dermis and epidermis, and in skin-exposed and skin-protected regions, for 100 subjects from a wide range of skin types (Fitzpatrick types I-V) and ages (13-83). For the statistical analysis, we report whether comparisons of the dermis and epidermis and sun-exposed and sun-protected areas across age and skin type are statistically significant, indeterminate, or not statistically significant and present 95% confidence intervals for this parameter as it ranges across different ages and skin types. This process of pre-analyzing features using healthy images provides a roadmap for how to ease the recruitment process while acquiring a sufficient range of images for effective AI model-building. We expect this type of analysis can have the effect of accelerating translation of AI-based OCT image analysis to the clinic.
Though patterns of anatomic distribution and morphology have been studied in a cross-sectional cohort of patients with skin chronic graft-versus-host disease (cGVHD; 1), their evolution within individual patients has not been well-characterized. Moreover, the value of medical photography for monitoring cGVHD remains unclear. In this pilot study, our objective was to retrospectively evaluate progression of cutaneous cGVHD over time using high-quality medical photographs and standardized avatar mapping, and correlate with subjective disease activity (2). After Institutional Review Board approval, we included 5 adult patients with cutaneous cGVHD who had the greatest number of serial photography sets over time. From photographs taken from onset of acute GVHD to last follow-up, combinations of epidermal, sclerotic and pigmentary morphologies were painted on a customized anatomy mapping application [AnatomyMapper]. Clinician and patient subjective impressions regarding disease severity at timepoints correlating with photographs were also documented. All 5 patients had Fitzpatrick type I-II skin and had undergone HLA-matched peripheral blood stem cell transplantation for hematologic malignancies. Generally, disease began as erythema on the trunk and upper extremities and progressed to pigmentary and sclerotic changes in the areas affected by erythema. Sclerotic isomorphic response was noted at the waistline. Visual changes on standardized avatars correlated subjectively with patient and clinician impression. The use of standardized avatar annotation of disease from medical photographs enables visualization of chronological progression of cGVHD morphology and subjectively correlates with clinician and patient impression of disease activity. Further studies are needed to determine consistency of disease patterns over a larger cohort and prognostic implications.