Human interaction plays an important role in medical image analysis, involving tasks such as manual labeling, editing, comparison, and assessment of image data. In addition, modern foundational models often rely on user input to define tasks and generate predictions, underscoring the importance of human involvement. Interacting with medical images requires visualization support for efficient rendering and navigation of 3D images, which can be challenging and time-consuming to build from scratch. This work introduces a Python library designed to simplify the creation of interactive interfaces for medical image analysis. It offers visualization and basic interaction for volumetric images, binary masks, 3D meshes, and videos, with additional user-defined interactions well supported. The library allows for highly customizable interfaces without limitations on view layouts or the number of images displayed simultaneously, unlike existing tools like 3DSlicer, ITK-SNAP, and Napari. Additionally, its Python-based design ensures compatibility with most deep-learning methods, reducing development and operational costs compared to heterogeneous visualization solutions. We demonstrate low interactive latency with our library and showcase its versatility through six interactive applications covering diverse practical scenarios. Our code is open-sourced at https://github.com/MIP-Lab/pytk-snap.git.
OBJECTIVES:Deep brain stimulation (DBS) of the globus pallidus internus (GPi) can reduce levodopa-induced dyskinesias, yet GPi targeting sometimes fails or worsens them, suggesting distinct prokinetic/antikinetic circuit mechanisms. Using the broad globus pallidus (GP) sampling from the Veteran Affairs Cooperative Study Program 468 (CSP-468) trial, we sought to identify stimulation zones and connectivity maps that predict dyskinesia outcomes in patients treated with GPi-DBS. MATERIALS AND METHODS:Data were drawn from the multicenter, randomized CSP-468 study which compared subthalamic nucleus with GPi-DBS in Parkinson disease. Postoperative imaging from patients treated with bilateral GPi-DBS with substantial presurgical dyskinesia (n = 69) was analyzed to relate lead location to change in hours of dyskinesia at six months postop. Cranial Suite and LeadDBSv3.0 enabled nonlinear registration to Montreal Neurologic Institute space and generation of volumes-of-tissue-activated (VTAs). Patients were pseudorandomly divided into model or hold-out cohorts. Outcome-weighted VTAs in the model cohort were used to derive above-average (sweet spot) and below-average (suboptimal) stimulation zones through t-tests. Outcome-weighted structural fibers were identified using the Parkinson Progressive Marker Initiative connectome. Normative resting state functional magnetic resonance imaging (rsfMRI) connectivity maps were generated using patient VTAs and the Brain Genomics Superstruct Project-1000 connectome, visualizing regions with significant synchrony differences between patients with improved vs worsened dyskinesias. All models were tested on the hold-out cohort. RESULTS:The sweet spot/suboptimal zone predicted outcomes in the hold-out cohort (RSpearman = 0.55, p = 0.004, q = 0.018). The sweet spot localized to the ventral posterior GPi, extending below the GPi, whereas the suboptimal zone capped the GPi UPDRS sweet spot, spanning mid-medial GP externa to dorsal GPi to mid-medial GPi. Neither overlapped the Unified Parkinson's Disease Rating Scale sweet spot. Fibers linked to antidyskinetic effects coursed beneath the GPi, similar to the path of the Ansa Lenticularis (RSpearman = 0.46, p = 0.02, q = 0.028). Normative rsfMRI maps showed greater supplementary motor area or premotor synchrony correlated with worse outcomes (RSpearman = 0.5, p = 0.008, q = 0.018). CONCLUSIONS:This analysis identified stimulation zones and connectivity that predict dyskinesia outcomes in an independent cohort, underscoring the relevance of the (sub)ventral GPi, Ansa Lenticularis, and supplementary motor area or premotor networks.
Dermatologic imaging has been rapidly expanding, with over 70% of related PubMed articles published since 2016 and over a million images across international research challenges and large-scale datasets with skin images. To improve data quality and usability, standardizing dermatologic imaging data for non-protected health information (non-PHI) research systems is essential. While the International Skin Imaging Collaboration (ISIC) has advanced standards in skin imaging, the field lacks a generalizable infrastructure to organize and describe imaging data for non-PHI research systems. This results in inconsistently labeled, heterogeneous datasets that hinder data integration, scalability, and interoperability. To address this gap, we propose the Dermatology Imaging Data Structure (DermIDS), inspired by the Brain Imaging Data Structure (BIDS) for neuroimaging. This structured framework aims to improve usability across datasets, reveal metadata gaps, and enable scalable artificial intelligence (AI)/machine learning (ML)-ready workflows. To illustrate this system, we curated and processed 1,000,692 images with DermIDS. We demonstrate that DermIDS (1) supports multimodal photographic data acquired from clinical photography, general photography, dermoscopy, reflectance confocal microscopy, and surface 3D imaging; (2) facilitates image-specific technical and clinical metadata organization; and (3) streamlines quality control and harmonization. Across all images, 1,256 unique metadata features were identified. However, 70% of clinical metadata features and 98% of technical metadata features were present in less than 100k images, highlighting key gaps and demonstrating the utility of DermIDS in revealing inconsistencies and opportunities for standardization. Our work supports large-scale analysis and harmonization, laying the foundation for AI/ML-ready workflows to advance dermatologic imaging research.
Transcranial-focused ultrasound (FUS) is a non-invasive neuromodulation technique capable of targeting deep brain regions with high precision. However, its mechanisms of action - particularly how it modulates neuronal activity and relates to fMRI BOLD signals - remain poorly understood. Using the macaque monkey thalamus and insular cortex as a model system, we show that low-intensity FUS evokes localized BOLD increases and preferentially modulates low-frequency (Delta, Theta, and Alpha) local field potentials and single-unit spiking in the ventroposterior lateral (VPL) nucleus, as well as through remote stimulation of the insular cortex. Compared with vibrotactile stimulation, FUS-evoked responses exhibit slower peak latencies. Importantly, we observe strong spatial correspondence between ultrasound-induced changes in spiking, LFP activity, and BOLD signals. These findings define the temporal and spatial characteristics of FUS neuromodulation and demonstrate the feasibility of precise, non-invasive targeting of deep brain circuits, supporting its translational potential for neuroscience and therapeutic applications.
We propose a novel method for establishing correspondence between two sequences of 2D images. One particular application of this technique is slice-level content navigation, where the goal is to localize specific 2D slices within a 3D volume or determine the anatomical coverage of a 3D scan based on its 2D slices. This serves as an important preprocessing step for various diagnostic tasks, as well as for automatic registration and segmentation pipelines. Our approach builds sequence correspondence by training a network to learn how to insert a slice from one sequence into the appropriate position in another. This is achieved by encoding contextual representations of each slice and modeling the insertion process using a slice-to-slice attention mechanism. We apply this method to localize manually labeled key slices in body CT scans and compare its performance to the current state-of-the-art alternative known as body part regression, which predicts anatomical position scores for individual slices. Unlike body part regression, which treats each slice independently, our method leverages contextual information from the entire sequence. Experimental results show that the insertion network reduces slice localization errors in supervised settings from 8.4 mm to 5.4 mm, demonstrating a substantial improvement in accuracy.
Skin segmentation from clinical photography is a crucial step in dermatological image analysis. However, the variability in skin tones, lighting conditions, anatomical regions, and the presence of additional objects introduces significant challenges. Due to these complexities, the segmentation process is often performed manually, as developing an algorithm capable of handling such diverse conditions is particularly difficult. Recently, openworld foundation models have emerged, offering the potential to generalize across diverse and unseen conditions. These models present a promising opportunity for dermatology. In this work, we adopt two such models-Grounding DINO and SAM 2-to construct a pipeline for zero-shot skin segmentation in dermatology. We evaluated our approach on two clinical skin photography datasets comprising 27,378 images. Based on a manual rating protocol, 77.1% of the segmentations were deemed acceptable, demonstrating robustness in handling realworld clinical photographs. Our results highlight the potential of open-world foundation models to address a challenging problem in dermatology with minimal human involvement.
Computed tomography (CT) images obtained in clinical settings are often acquired with diverse scanner types and acquisition parameters. They may exhibit significant variations in fields of view (FOVs) and levels of contrast enhancement. An automated method for navigating the content in these images is therefore essential for effective dataset curation and downstream analyses. This work introduces a framework called Body- Part-Phase Regression (BPPR) to automatically identify regions of interest and determine the contrast enhancement phase of body CT images. The framework consists of two key components: (1) A two-phase body part regression method for predicting the anatomical location of 2D slices within 3D volumes. (2) A circular regression model for predicting the contrast timing of CT images (i.e., the timing of the scan relative to contrast agent injection) from a continuous perspective, providing a fine-grained understanding of contrast differences, particularly in relation to patient-specific vascular effects. These two components are linked via a positional weighting mechanism which enhances volumelevel phase prediction by leveraging slice-level predictions. By unifying the “part” and “phase” regression models, our framework establishes a cohesive approach to continuous content navigation in CT images. We train and evaluate our models on large-scale datasets consisting of multi-contrast images and compare their performance with alternative approaches pursuing similar goals. The experiments demonstrate improvements in both slice localization and contrast phase prediction. In particular, the two-phase training scheme reduces the slice localization error of previous body part regression methods from 9.2 mm to 6.1 mm. We also discuss the distinctive advantages of BPR over segmentationbased approaches and highlight potential clinical applications that may benefit from the proposed BPPR framework.
We designed and successfully implemented photography protocols for the PALM007 (Pamoja Tulinde Maisha - Together Save Lives in Swahili) randomised clinical trial evaluating tecovirimat for the treatment of mpox in the Democratic Republic of the Congo (DRC). We addressed the unique challenges and limitations of conducting standardised photography procedures on patients with mpox in high-risk 'red zones' in remote study treatment centre sites. Key considerations included standardised photography protocols, training, quality assurance, participant privacy and comfort, and interpersonal communication. Our developed procedures enabled acquisition of 61 926 standardised, clinical-quality photographs of 597 patients with mpox. These considerations may guide future studies collecting standardised photographs in a low-resource setting.
The cross-Modality Domain Adaptation (crossMoDA) challenge series, initiated in 2021 in conjunction with the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), focuses on unsupervised cross-modality segmentation, learning from contrast-enhanced T1 (ceT1) and transferring to T2 MRI. The task is an extreme example of domain shift chosen to serve as a meaningful and illustrative benchmark. From a clinical application perspective, it aims to automate Vestibular Schwannoma (VS) and cochlea segmentation on T2 scans for more cost-effective VS management. The challenge has evolved across three editions to address increasingly complex clinical scenarios: beginning with single-institution controlled data for binary segmentation (tumour and cochlea) in 2021, extending to multi-institution data with Koos grade classification in 2022, and culminating in heterogeneous routine surveillance data with intra- and extra-meatal tumour sub-segmentation in 2023. In this work, we report the findings of the 2022 and 2023 crossMoDA editions alongside a retrospective analysis of challenge progression. Successive editions demonstrate a reduction in segmentation outliers despite concurrently increasing dataset heterogeneity, suggesting that expanded training diversity improves model robustness. Notably, the 2023 winning approach generalised effectively to prior homogeneous test sets, evidencing the benefit of heterogeneous training data. However, cochlear Dice scores declined in 2023. This decrease may reflect the combined effects of increased dataset heterogeneity, resolution variability, and the added optimisation complexity of a three-class segmentation task, although the precise cause remains uncertain. Progress can and should still be made on the specific task of VS segmentation to reach clinical acceptability but as the performance of leading competitors starts to plateau, a more challenging cross-modal learning task may be beneficial to serve as a benchmarking tool in the future.
Deep brain stimulation (DBS) of the globus pallidus interna (GPi) is an effective treatment for medication‑refractory Parkinson's disease (PD), though outcomes vary widely. The influence of lead location relative to the GPi therapeutic sweetspot has not previously been examined within a multivariable regression framework. This study evaluated the predictive utility of volume-of-tissue-activated (VTA)-sweetspot overlap on postoperative UPDRS‑III outcomes in the CSP #468 cohort receiving bilateral GPi‑DBS, alongside demographic/clinical predictors, using two independently derived sweetspot models. CSP #468 was a multicenter randomized clinical trial with blinded 6‑month outcomes and robust collection of associated features of PD. A prior publication generated a cross validated 6‑month GPi-sweetspot from this dataset. An independent single‑surgeon cohort was used to construct a separate sweetspot uninfluenced by CSP outcomes. Multivariable modeling was completed using backward‑selection linear regression with Bonferroni correction. Models were checked for assumptions, data leakage, and overfitting. Each sweetspot-multivariable model's ability to predict outcomes in the opposite cohort was assessed. Both sweetspots localized to the primary motor GPi. The independent sweetspot multivariable model included VTA‑sweetspot overlap and levodopa response (R²Adj = 0.19) and remained robust under leave-one-out cross validation. The CSP sweetspot multivariable model also predicted outcomes in the independent cohort (p = 0.004), demonstrating meaningful external validity.
Purpose:Clinical photographs play an integral role across medical fields. Since the mid-20th century, deidentification has consisted of black bars covering specific facial features, typically the eyes alone. Although increasingly questioned, this practice persists in clinical and academic settings. Approach:A barrier to standardized deidentification guideline development is the unknown risk of artificial intelligence (AI) to reconstruct faces from partially obscured photos. We evaluate the ability of generative AI to reconstruct 10,000 facial images in the Synthetic Faces High Quality dataset across 14 regional masking strategies. Results:Covering the eyes or any other single facial feature resulted in highly identifiable reconstructions, demonstrated by low face mesh distortion (0.14 to 0.18 relative to whole-face masking; absolute total face mesh distortion 8.34 to 10.19) and high structural similarity index to the original face (1.24 to 1.25 relative to whole-face masking; absolute SSIM 0.91 to 0.92). An open-source face verification model using Dlib was able to match 97.98% to 99.93% of these reconstructed images with the original image prior to single feature masking. Removing all major facial features (eyebrows, eyes, nose, and mouth) resulted in a threefold reduction in face verification rates compared with eyes alone, from 98.87% (95% CI [98.63%, 99.07%]) to 33.93% (95% CI [32.95%, 34.94%]). Conclusions:We provide quantitative metrics of the reidentification risk that modern generative AI technology poses for partially obscured facial images.
Standardized 2D photography plays an essential role in dermatologic practice, supporting longitudinal documentation, patient monitoring, and consensus-based clinical scoring. However, photographs taken from limited views often suffer from reduced anatomical context and missing body-location information. 3D representations enable unified spatial interpretation of multi-view imagery. Recent developments in computer vision have made it feasible to infer dense correspondences between 2D images and a 3D human mesh. In this study, we explored integrating 2D dermatological images with a 3D surface model using DensePose, a deep learning-based human dense correspondence framework. This creates an anatomically grounded representation that supports mesh-level analyses and recovers spatial context for each image. We use a dataset including four full body photographs (front, back, and each side) from each of 147 subjects with chronic graft-versus-host disease, for a total of 588 images. Our method integrates these multiple 2D full-body photographs captured across varied body shapes and camera angles into a 3D mesh. We further showed that the resulting 3D mesh enables quantification of the extent to which individual 2D images, or their combinations, represent the complete body surface. On average, a single full body view captures 28% of the body surface, while adding a second, third, and fourth view increases average coverage to 50%, 72%, and 80%, respectively. To assess spatial consistency, we annotated up to 10 anatomical landmarks per patient on 80 images across 20 patients and reported a median pairwise geodesic distance between corresponding landmarks of 4.6 cm. These findings can guide how dermatology images are captured and support future opportunities in monitoring, education, and communication using existing infrastructure.
OBJECTIVE:Essential tremor (ET) is a prevalent movement disorder that also includes nonmotor symptoms such as anxiety, depression, and cognitive impairment. Deep brain stimulation (DBS) is an established treatment for ET, yet its impact on nonmotor symptoms remains unclear. This study aims to describe neuropsychological outcomes following ventral intermediate nucleus (VIM) DBS in a large cohort of patients with ET and identify factors associated with changes in depression and cognitive function. METHODS:A retrospective cohort study of patients who had undergone VIM DBS was performed. Inclusion criteria were ET diagnosis, surgery between October 2007 and March 2020, and available pre- and post-DBS neuropsychological testing results. Neuropsychological measures included the Beck Depression Inventory-II (BDI-II), Beck Anxiety Inventory (BAI), and cognitive measures assessing attention, executive function, language, memory, and visuospatial function. Post-DBS tremor improvement was graded, and active electrode coordinates and stimulation parameters were identified. Statistical analyses included descriptive statistics, t-tests to compare pre- and postoperative scores at the group level, and one-way analysis of variance to compare variables among patients who improved, were stable, or worsened in psychiatric and cognitive characteristics after DBS. RESULTS:One hundred thirty-nine patients met the study inclusion criteria. BDI-II scores significantly decreased postoperatively (9.82 ± 6.77 vs 8.29 ± 6.18, p < 0.001, Cohen's d = 0.176), whereas BAI scores remained unchanged. Both language (p = 0.003, Cohen's d = 0.259) and memory (p < 0.001, Cohen's d = 0.336) domains showed statistically significant small-magnitude declines following surgery, whereas attention, executive function, and visuospatial function were unchanged. Patients with improved depression (14.3%) following VIM DBS had significantly higher BDI-II scores preoperatively (p < 0.001, ω2 = 0.226). Patients with worsened language (18.7%) had higher preoperative language scores (p < 0.001, ω2 = 0.058). Patients with worsened memory (15.1%) had higher BAI scores preoperatively (p = 0.002, ω2 = 0.079). Preoperative scores were similar between patients with improved and worsened overall cognition postsurgery. Patients with improved overall cognition had improvements in attention, language, and visuospatial function. CONCLUSIONS:VIM DBS for ET did not result in large-magnitude neuropsychological changes. There were statistically significant, though likely not clinically meaningful, small-magnitude improvements in depression and worsening in language and memory scores. Associations were found between multiple preoperative mood and cognitive scores and post-DBS neuropsychological changes. These findings can help inform clinical decision-making and patient counseling for DBS.
Localization of ear landmarks in full head computed tomographic (CT) images can facilitate automated analysis of the inner ear anatomy by identifying and aligning the regions-of-interest (ROIs) around the ear. This study explores unsupervised methods for localizing ear landmarks using the Self-Supervised Anatomical Embeddings (SAME) framework, which leverages contrastive learning to learn voxel-wise anatomical representations for 3D volumes. We define seven landmarks in an atlas image and localize these landmarks in testing images by comparing the embeddings of the atlas landmarks to the embeddings of each voxel in the testing images. Our experiments show that this unsupervised method achieves results comparable to a convolutional neural network trained to predict the same landmarks in a supervised manner. Importantly, the unsupervised approach does not require ground-truth labels for training and can be applied to localize other landmarks without retraining. Additionally, we explore metrics to assess the reliability of landmarks predicted using the unsupervised method. We find that the number and distribution of voxels in the test image with high embedding similarity to a given landmark in the atlas serves as a good indicator of the expected localization error. This insight allows for the selection of points that can be more reliably localized without the need for direct measurement of localization errors, which would require ground-truth labels. This study demonstrates the potential of unsupervised learning techniques for anatomical landmark localization in head CT images, offering a scalable solution that eliminates the need for extensive manual annotation.
Introduction Benign and malignant myxoid soft tissue tumors have shared clinical, imaging, and histologic features that can make diagnosis challenging. The purpose of this study is comparison of the diagnostic performance of a radiomic based machine learning (ML) model to musculoskeletal radiologists. Methods Manual segmentation of 90 myxoid soft tissue tumors (45 myxomas and 45 myxofibrosarcomas) was performed on axial T1, and T2FS or STIR magnetic resonance imaging sequences. Eighty-seven radiomic features from each modality were extracted. Five ML models were trained to classify tumors as benign or malignant in 40 tumors and then tested with an additional 50 tumors using cross validation. The accuracy of the best ML model based on area under the receiver operating characteristic curve (AUC) was compared to the consensus diagnosis of three musculoskeletal radiologists. Correlation between radiologist confidence (equivocal, probably, consistent with) and accuracy was tested. Results The best ML classifier was a logistic regression model (AUC 0.792). Using T1 + T2/STIR images, the ML model classified 78% (39/50) of tumors correctly at a similar rate compared to 74% (37/50) by radiologists. When radiologists disagreed, the consensus diagnosis classified 50% of tumors (7/14) correctly compared to 86% (12/14) by the ML model, though this did not reach statistical significance. Radiologists had a cumulative accuracy of 91% (30/33) when they rated their confidence ‘consistent with’ compared to 61% (31/51) when they rated their confidence ‘equivocal/probably’ (P = 0.006). For cases when radiologists rated their confidence ‘equivocal/probably’, the ML model had 76% accuracy (39/51). Conclusions A radiomic based ML model predicted benign or malignant diagnosis in myxoid soft tissue tumors similarly to the consensus diagnosis by three musculoskeletal radiologists. Radiologist confidence in the diagnosis strongly correlated with their diagnostic accuracy. Though radiomics and radiologists perform similarly overall, radiomics may provide novel diagnostic utility when radiologist confidence is low, or when radiologists disagree.
Contrast enhancement is widely used in computed tomogra-phy (CT) scans, where radiocontrast agents circulate through the bloodstream and accumulate in the vasculature, creating visual contrast between blood vessels and surrounding tissues. This work introduces a technique to predict the timing of contrast in a CT scan, a key factor influencing the contrast effect, using circular regression models. Specifically, we represent the contrast timing as unit vectors on a circle and employ 2D convolutional neural networks to predict it based on predefined anchor time points. Unlike previous methods that treat contrast timing as discrete phases, our approach is the first method that views it as a continuous variable, of-fering a more fine-grained understanding of contrast differences, particularly in relation to patient-specific vascular effects. We train the model on 877 CT scans and test it on 112 scans from different subjects, achieving a classification accuracy of 93.8%, which is similar to state-of-the-art results re-ported in the literature. We compare our method to other 2D and 3D classification-based approaches, demonstrating that our regression model have overall better performance than the classification models. Additionally, we explore the relation-ship between contrast timing and the anatomical positions of CT slices, aiming to leverage positional information to improve the prediction accuracy, which is a promising direction that has not been studied.
This prospective study investigated the potential benefits of deactivating the second most apical electrode to improve access to lower-frequency pitch and first formant information to help improve speech and music outcomes with a cochlear implant. Twenty-one adults (30 ears) with cochlear implants completed an A-B-A-B study to compare the participant's clinical map with all electrodes active (A) and their clinical map with the second most apical electrode deactivated (B). Test measures included pitch discrimination, speech understanding in noise, and subjective musical sound quality and enjoyment ratings. This study also investigated the impact of participant demographic and electrode placement factors on the degree of benefit derived from the experimental map (B). There was no significant difference between the two conditions on any measure at the group level. However, individual participants demonstrated improvements in pitch discrimination (33.3%), speech perception in noise (43.3%), musical sound quality (50.0%), and musical enjoyment (40.0%). Musical sound quality and enjoyment ratings were strongly correlated, and speech perception correlated with musical enjoyment but not sound quality. Electrodes outside scala tympani, smaller electrode-to-modiolus distances, and certain device manufacturers (Cochlear and MED-EL) predicted greater benefit from deactivating the second-most apical electrode. Certain adult cochlear implant users may benefit from selective apical electrode deactivation, depending on their demographic and electrode placement profile. Clinicians could consider deactivating the second most apical electrode with patients, who report poor musical sound quality or those who have disengaged from music since receiving their CI to assess potential benefits individually.
Mpox is a viral illness with heavy cutaneous involvement. Automatic tracking of mpox lesion progression is critical in determining the resolution of evolving lesions. This work introduces a novel application of deep learning for lesion monitoring through alignment of dermatological hand photographs. By adapting the VoxelMorph framework for 2D photographic data, we explore key point alignment across serial images. We trained our neural network model on a unique dataset of 1,658 hand images and evaluated its performance on a test set of 254 images. Additionally, we validated the method's generalizability with a supplementary set of 500 images, which included extensive Mpox infection. Our findings indicate modest yet significant improvements in key points and lesion center registration across different regularization strengths. Although promising, the complexity of hand structure presents challenges, requiring cautious application and further refinement, especially in regions with intense spatial discontinuities, such as interdigital areas.
Measuring skin involvement in chronic graft-versus-host disease (cGVHD) currently requires expert manual assessment, which is costly, time-consuming, and shows high interrater disagreement (>20% surface area). In our previous work, automated image analysis showed promise for measuring affected skin area under controlled photography conditions. Our aim is to improve the performance of these methods in standard clinical photographs without the need for costly expert annotations using a semi-supervised approach. A baseline U-Net model was trained in a fully supervised manner using 360 3D photographs from 36 cGVHD patients, with expert-marked ground truth contours of affected skin. The model was then iteratively retrained by incorporating an additional 5648 unlabeled photographs from 83 new patients using a semi-supervised method. Testing on clinical photographs of 20 held-out patients, the median surface area error improved from 19.2% (interquartile range 6.3 - 33.8) at baseline to 10.2% (4.5 - 22.6) after retraining. Semi-supervised training therefore provides an effective method for translating a pre-trained U-Net segmentation model to standard clinical photographs, without the need for additional expert annotations. Such models could help standardize cGVHD assessment and tracking, alleviating the need for costly expert evaluations and providing a reliable tool that would significantly enhance the current standard of manual assessment.