Abstract Background The teaching of bilateral sagittal split osteotomy (BSSO) requires the utilisation of practical training methods to optimise the surgical result and prevent complications such as bad splits or damage to the inferior alveolar nerve. The objective of the present study was to evaluate a structured training program using a cost-effective, 3D-printed mandibular model to improve osteotomy and splitting performance in BSSO. Material & methods The study was registered in the German Clinical Trials Register (DRKS00034369, registration date: 20/11/2024). A lower jaw model was fabricated using a semi-professional Prusa XL Multi-Tool 3D printer and a combination of colours and materials. The model also incorporated the inferior alveolar nerve, which was fabricated from a flexible material. The model was modified to facilitate its dynamic integration into a phantom head by modification of the condyles. Twenty participants performed two repetitions of the osteotomy and splitting steps of BSSO. The participants received feedback after the first run and during the second run. The experiments were recorded and assessed using a validated questionnaire (OSATS). Furthermore, the distance between the inferior alveolar nerve and the fracture surface, as well as the number of bad splits, was determined. Results Participants significantly improved their surgical skills as measured by the OSATS between the first and second run from 23.7 points ± 4.3 (mean ± sd) to 27.2 points ± 3.9 (p < 0.001). In addition, the minimal distance between the nerve and the fracture surface increased from 0.2 ± 0.38 mm to 1.12 ± 1.04 mm (p = 0.012). The number of bad splits decreased from 15 to 4 (p = 0.044). The estimated material cost of the lower jaw is approximately US$4 with a printing time of an estimated 9.5 h. Conclusion It is possible to develop a cost-effective and realistic training model, that can be used within structured repeated practice sessions for BSSO training and is associated with improved performance in this setting. Further studies are needed to validate the simulator and assess its impact on real-life surgical outcomes.
Diffusion models produce high-quality synthetic data but suffer from slow inference. We propose 3D Variable-Step Denoising Diffusion Probabilistic Model (VS-DDPM) a framework engineered to maintain generative quality while accelerating inference by several factors. We tested our approach on four tasks (missing MRI, tumor removal, MRI-to-sCT, and CBCT-to-sCT) within the BraTS2025 and SynthRAD2025 challenges. Designed for high efficiency under hardware and time constrains imposed by both challenges. VS-DDPM achieved state-of-the-art (SOTA) performance in missing MRI synthesis, yielding Dice scores of 0.80, 0.83, and 0.88 for the enhancing tumor, tumor core, and whole tumor regions, respectively, alongside a structural similarity index (SSIM) of 0.95. For MRI tumor removal, the model attained a root mean squared error (RMSE) of 0.053, a peak signal-to-noise ratio (PSNR) of 26.77, and an SSIM of 0.918. While the framework demonstrated competitive performance in MRI-to-sCT and CBCT-to-sCT tasks, it did not reach SOTA benchmarks, potentially due to sensitivities in data pre and post-processing pipelines or specific loss function configurations. These results demonstrate that VS-DDPM provides a robust and tunable solution for high-fidelity 3D medical image synthesis. The code is available in https://github.com/andre-fs-ferreira/SynthRAD_by_Faking_it.
IntroductionAs one of the first research teams with full access to Siemens’ Cinematic Reality, we evaluated its usability and clinical potential for cinematic volume rendering on the Apple Vision Pro.MethodsWe visualized venous-phase liver computed tomography and magnetic resonance cholangiopancreatography scans from the CHAOS and MRCP_DLRecon public datasets, respectively. Fourteen medical experts assessed usability and anticipated clinical integration potential using standardized questionnaires (System Usability Scale and ISONORM 9242-110-S) and an open-ended survey. Our primary aim was not to validate direct clinical outcomes, but to evaluate the usability of Siemens’ Cinematic Reality on the Apple Vision Pro and gather expert feedback on potential use cases and missing features required for clinical adoption beyond educational purposes.ResultsTheir feedback identified feasibility, key usability strengths, and required features to catalyze the adaptation in real-world clinical workflows.ConclusionThe findings provide insights into the potential of immersive cinematic rendering in medical imaging and the needed features for clinical adoption as suggested by the medical experts. Siemens Cinematic Reality running on the Apple Vision Pro was deemed to have good usability, making it a promising tool.
Abstract Adverse drug effects remain a major barrier to safe and effective cancer therapy, underscoring the need for tools that predict treatment-related toxicities. We analyzed multimodal real-world data from 14,596 cancer patients across 38 cancer entities, encompassing 330 clinical, tumor, and imaging characteristics, along with 89 anticancer agents. Hematological adverse events (HAE), defined by nadirs of hemoglobin, leukocyte, neutrophil, and platelet values within two months of treatment initiation, were highly prevalent (87.7%; 33.1% severe). We developed Toxix , an explainable artificial intelligence (xAI) framework modeling interactions between patient characteristics and drug combinations. Toxix achieved strong predictive performance for severe toxicities (median AUROC 0.85 for anemia; >0.76 for leukopenia, neutropenia, and thrombocytopenia) and was validated in an external cohort of 2,768 patients with non-small cell lung cancer. Model explainability enabled systematic characterization of drug-patient interactions underlying HAEs. Toxix provides a real-world informed framework for personalized and toxicity-aware cancer therapy planning.
Background:Augmented reality head-mounted displays could overcome the spatial dissociation between medical imaging and the surgical field, which may be particularly important in anatomically dense regions, such as the head and neck. Although many head-mounted displays offer markerless inside-out tracking at a fraction of the cost of navigation systems, their overlay accuracy with superimposition (SI) modality onto the surgical field remains limited. The virtual twin (VT), displaying holography adjacent to the surgical field, may offer a viable alternative. However, its performance is still unclear. Objective:This study aimed to compare the accuracy and efficiency of the two visualization modalities, SI and VT, for anatomical localization in the head and neck region. Methods:In a randomized crossover trial to compare two augmented reality visualization modalities (SI and VT), 38 participants used a HoloLens 2 to localize point, line-based, and volume-based anatomical structures on head phantoms. Their performance was evaluated with respect to accuracy, workload, time, and user experience. Results:SI achieved significantly better point localization accuracy than VT both in absolute (mean 14.4, SD 4.2 mm vs mean 15.8, SD 5.5 mm; P=.003) and relative accuracy (mean 3.4, SD 2.2 mm vs mean 6.0, SD 5.0 mm; P<.001). In line-based structures, accuracy was comparable between SI (average surface distance [ASD], mean 23.4, SD 4.1 mm; Hausdorff distance [HD], mean 31.5, SD 7.8 mm) and VT (ASD=mean 23.0, SD 4.5 mm; P=.51; HD=mean 31.0, SD 7.5 mm; P=.57). However, SI showed significantly higher deviation than VT in volume-based structure (ASD=mean 37.1, SD 13.8 mm vs mean 34.1, SD 14.2 mm; P=.01; HD=mean 52.0, SD 16.8 mm vs mean 49.1, SD 15.8 mm; P=.03). Participants were faster with SI (P=.02), while workload NASA-TLX (National Aeronautics and Space Administration Task Load Index) scores did not demonstrate a significant difference (P=.79). Conclusions:Given that SI did not clearly outperform VT under overlaid soft tissue and viewing challenges, VT remains a viable alternative in certain surgical scenarios where high accuracy is not required. Future research should focus on optimizing viewing angle guidance and the linkage between the anatomical target and the skin surface.
Due to significant research efforts in the field of virtual reality (VR) and the increasing interest of the gaming industry, technical possibilities have developed rapidly. VR applications have become more extensive in recent years, with devices becoming more convenient and less complex to handle. The purpose of this review was to summarize the current state of research on the use of VR in oncology and to discuss its potential role in patient-centered care. As a result of technological progress, the benefits of VR have entered the clinical field and are being used in various areas. However, since VR has not yet become part of everyday life, most patients are inexperienced in using these devices. Previous studies have shown that VR applications for oncology patients can improve well-being and reduce stress, anxiety, and nausea during therapy, surgery, and diagnostic procedures. Oncology is particularly well suited for psychological VR applications because cancer can be a potentially traumatic experience with a strong impact on mental health, which in turn can influence therapeutic success. In addition, oncological treatments are usually long-term, making them well suited for studies involving extended VR use. VR represents a promising supportive tool in oncology, particularly for improving psychological well-being and enhancing patient-centered care. Ongoing research suggests that its use may positively influence both the mental state of patients and the overall treatment experience.
Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment holds significant potential to improve surgical training. While machine learning-based methods are increasingly popular for assessing skills in minimally invasive surgery, their application to open surgery remains limited. We present the results of a dedicated MICCAI challenge designed to benchmark and advance vision-based skill assessment in open surgery. The challenge dataset comprises videos of an open suturing training task recorded with a static GoPro camera in a dry-lab setting, with instrument trajectories available in addition to the primary video modality. The OSS Challenge was hosted over two consecutive years, comprising two and three independent tasks, respectively: (1) classifying skill level into four classes, (2) predicting the full Objective Structured Assessment of Technical Skills across eight categories, and (3) tracking hands and surgical tools. Participants submitted diverse solutions including deep learning-based video models, tracking-driven methods, and hybrid approaches. General-purpose spatiotemporal video models consistently achieved the strongest performance, though conceptually diverse approaches reached competitive levels when well-executed. Predicting fine-grained OSATS scores remains challenging but benefits substantially from increased training data. Keypoint tracking proves difficult given frequent occlusions and out-of-frame instances, limiting current applicability for motion-based skill analysis. This work benchmarks innovative and diverse solutions for surgical skill assessment, highlighting both the promise and current limitations of video-based evaluation in open surgery and identifying critical directions for advancing automated skill assessment toward clinical impact.
Accurate multi-class skin lesion classification from dermatological images is critical for effective patient care, though often hindered by limited harmonized data and unaddressed skin tone biases. This work presents an analysis of multi-class classification and bias quantification using a large, harmonized image dataset compiled from 13 public sources. We developed and validated an effective deep learning model for skin tone classification from images, which then pseudo-labeled unannotated images, revealing significant imbalances favoring lighter skin tones. Our multi-class models demonstrated strong overall performance in classifying six common skin lesion types. Critically, bias analysis revealed consistently higher model performance metrics on darker skin tones compared to lighter tones, despite significant underrepresentation of darker tones in the data. This counterintuitive finding calls for dedicated research to understand these disparities and prioritize the acquisition of more balanced, human-validated datasets, essential for developing fair and reliable AI solutions for all patient populations.
Despite the advances in automated medical image segmentation, AI models still underperform in various clinical settings, posing challenges for integration into real-world workflows. In this pre-registered prospective multicenter evaluation, we analyzed 20 state-of-the-art mandibular segmentation models across 19,218 segmentations of 1,000 clinically resampled CT/CBCT scans. Our results suggest that for a given model, segmentation accuracy can vary by up to 25% in Dice score as socio-technical factors such as voxel size, bone orientation, and patient conditions (e.g., osteosynthesis or pathology) shift from favorable to adverse. Higher sharpness, isotropic smaller voxels, and neutral orientation significantly improved results, while metallic osteosynthesis and anatomical complexity led to significant degradation. Our findings challenge the common view of AI models as “plug-and-play” tools and suggest evidence-based optimization recommendations for both clinicians and developers. This will in turn boost the integration of AI segmentation tools in routine healthcare.
The nnU-Net has demonstrated continuous success in medical segmentation tasks, which heavily rely on the availability and diversity of annotated biomedical data. However, assembling medical imaging cohorts remains challenging due to numerous factors such as privacy regulations and annotation costs. As a result, data augmentation plays a crucial role in increasing data availability while maintaining anatomical feasibility. Hence, we propose the ++nnU-Net, a novel data augmentation module based on image registration that operates prior to preprocessing and training take place. Our framework was evaluated across five different 2D datasets. In this workflow, image data go through a two-stage registration process, generating new warped images. The transformations are then applied to the respective segmentation. In addition, the pipeline computes available disk space, generates supplementary binary synthetic masks and generates checkpoints. We demonstrate that the ++nnU-Net outperforms the nnU-Net baseline, yielding improvements in Dice Similarity Coefficient scores. In the most prominent cases, we observe performance gains of approximately 22%. These findings highlight the effectiveness of registration-based data augmentation, particularly for 2D medical imaging datasets and suggest that the ++nnU-Net provides a practical and scalable approach for enhancing segmentation performance in data-limited settings. The source code for the ++nnU-Net is available at: https://github.com/sofia-adelie/plusplusnnunet.git
Head and neck cancer (HNC) patients face an increased risk of malnutrition due to lifestyle, tumor localization, and treatment effects. While skeletal muscle area (SMA) and radiation attenuation (SM-RA) at the third lumbar vertebra (L3) are established prognostic markers, L3 is not routinely available in head and neck imaging. The prognostic value of SM-RA at the third cervical vertebra (C3) remains unclear. This study assesses whether SMA and SM-RA at C3 predict locoregional control (LRC) and overall survival (OS) in HNC. We analyzed 904 HNC cases with head and neck CT scans. A deep learning pipeline identified C3, and SMA/SM-RA were quantified via automated segmentation with manual verification. Cox proportional hazards models assessed associations with LRC and OS, adjusting for clinical factors. Median SMA and SM-RA were 36.64 cm² (IQR: 30.12–42.44) and 50.77 HU (IQR: 43.04–57.39). In multivariate analysis, lower SMA (HR 1.62, 95% CI: 1.02–2.58, p = 0.04), lower SM-RA (HR 1.89, 95% CI: 1.30–2.79, p < 0.001), and advanced T stage (HR 1.50, 95% CI: 1.06–2.12, p = 0.02) were prognostic for LRC. OS predictors included advanced T stage (HR 2.17, 95% CI: 1.64–2.87, p < 0.001), age ≥70 years (HR 1.40, 95% CI: 1.00–1.96, p = 0.05), male sex (HR 1.64, 95% CI: 1.02–2.63, p = 0.04), and lower SM-RA (HR 2.15, 95% CI: 1.56–2.96, p < 0.001). Deep learning-assisted SM-RA assessment at C3 outperforms SMA for LRC and OS in HNC, supporting its use as a routine biomarker and L3 alternative.
In this article, we present a brain tumor database collection comprising 23,049 samples, with each sample including four different types of MRI brain scans: FLAIR, T1, T1ce, and T2. Additionally, one or two segmentation masks (ground truth) are provided for each sample. The first mask is the raw output from the registration process and is provided for all samples, while the second mask, provided particularly for synthetic samples, is a post-processed version of the first, designed to simplify interpretation and optimize it for network training. These samples have been acquired via registration process of 438 samples available at the moment of registration from the original dataset provided by the BraTS 2022 Challenge. Registering each pair of existing brain scans results in two additional scans that retain a similar brain shape while featuring varying tumor locations. Consequently, by registering all possible pairs, a dataset originally consisting of n samples can be expanded to n2 samples. The original dataset was collected from different institutions under standard clinical conditions, but with different equipment and imaging protocols. As a result, the image quality is heterogeneous, reflecting the diversity of clinical practices across institutions. This dataset can be utilized for various tasks, such as developing fully automated segmentation algorithms for new, unseen brain tumor cases, particularly through deep learning-based approaches, since ground truth is provided for each sample.
Defects to human crania are one kind of head bone damages, and cranial implants can be used to repair the defected crania. The automation of the implant design process is crucial in reducing the corresponding therapy time. Taking the cranial implant design problem as a special kind of shape completion task, an automatic cranial implant design workflow is proposed, which consists of a deep neural network for the direct shape prediction of the missing part of the defective cranium and conventional post-processing steps to refine the automatically generated implant. To evaluate the proposed workflow, we employ cross-validation and report an average Dice Similarity Score and boundary Dice Similarity Score of 0.81 and 0.81, respectively. We also measure the surface distance error using the 95th quantile of the Hausdorff Distance, which yields an average of 3.01 mm. Comparison with the manual cranial implant design procedure also revealed the convenience of the proposed workflow. In addition, a plugin is developed for 3D Slicer, which implements the proposed automatic cranial implant design workflow and can facilitate the end-users.
Background In patients with acute vestibular syndrome (AVS) differentiating between benign acute peripheral vestibular disorders and possible life-threatening central, causes such as stroke, can be challenging due to similar symptoms. AVS patients experience dizziness, vertigo, imbalance, nausea, vomiting, and abnormal eye movements. This research evaluates the feasibility of using the eye-tracking capability of a mixed reality optical-see-through head-mounted display (MR-OST-HMD) to detect pathological eye movement patterns in patients with AVS.Methods Conducted at University Hospital Essen, this study assessed patients with AVS using a MR-OST-HMD during the HINTS-Exam. The feasibility study included 21 healthy subjects, seven patients with acute peripheral vestibular dysfunction and two stroke patients. Eye gaze, head position, and orientation were captured using a MR-OST-HMD and an in-house developed application designed to simulate the HINTS-Exam. The eye-tracking technology determined gaze direction and position, while the internal measurement unit and gyroscope recorded head movements in terms of position and velocity.Results The MR-OST-HMD detected abnormal eye movements, including nystagmus, saccades, and skew deviation effectively. The device proved effective even for patients with severe nausea and elderly participants, who completed the eye calibration and HINTS-Exam without difficulty. The MR-OST-HMD HINTS-Exam was quick to perform (approximately 5 min) and was easily integrated into clinical practice after a single demonstration for medical staff.Conclusion MR-OST-HMD can detect pathological eye movements in AVS patients. Future research should validate these findings in larger cohorts and explore machine learning integration to enhance diagnostic accuracy.
Shape reconstruction from imaging volumes is a recurring need in medical image analysis. Common workflows start with a segmentation step, followed by careful post-processing and,finally, ad hoc meshing algorithms. As this sequence can be timeconsuming, neural networks are trained to reconstruct shapes through template deformation. These networks deliver state-ofthe-art results without manual intervention, but, so far, they have primarily been evaluated on anatomical shapes with little topological variety between individuals. In contrast, other works favor learning implicit shape models, which have multiple benefits for meshing and visualization. Our work follows this direction by introducing deep medial voxels, a semi-implicit representation that faithfully approximates the topological skeleton from imaging volumes and eventually leads to shape reconstruction via convolution surfaces. Our reconstruction technique shows potential for both visualization and computer simulations.
Despite advances in precision oncology, clinical decision-making still relies on limited variables and expert knowledge. To address this limitation, we combined multimodal real-world data and explainable artificial intelligence (xAI) to introduce AI-derived (AID) markers for clinical decision support. We used xAI to decode the outcome of 15,726 patients across 38 solid cancer entities based on 350 markers, including clinical records, image-derived body compositions, and mutational tumor profiles. xAI determined the prognostic contribution of each clinical marker at the patient level and identified 114 key markers that accounted for 90% of the neural network's decision process. Moreover, xAI enabled us to uncover 1,373 prognostic interactions between markers. Our approach was validated in an independent cohort of 3,288 patients with lung cancer from a US nationwide electronic health record-derived database. These results show the potential of xAI to transform the assessment of clinical variables and enable personalized, data-driven cancer care.
The automated analysis of the aortic vessel tree (AVT) from computed tomography angiography (CTA) holds immense clinical potential, but its development has been impeded by a lack of shared, high-quality data. We launched the SEG.A. challenge to catalyze progress in this field by introducing a large, publicly available, multi-institutional dataset for AVT segmentation. The challenge benchmarked automated algorithms on a hidden test set, with subsequent optional tasks in surface meshing for computational simulations. Our findings reveal a clear convergence on deep learning methodologies, with 3D U-Net architectures dominating the top submissions. A key result was that an ensemble of the highest-ranking algorithms significantly outperformed individual models, highlighting the benefits of model fusion. Performance was strongly linked to algorithmic design, particularly the use of customized post-processing steps, and the characteristics of the training data. This initiative not only establishes a new performance benchmark but also provides a lasting resource to drive future innovation toward robust, clinically translatable tools.
It is an open secret that ImageNet is treated as the panacea of pretraining. Particularly in medical machine learning, models not trained from scratch are often finetuned based on ImageNet-pretrained models. We posit that pretraining on data from the domain of the downstream task should almost always be preferred instead. We leverage RadNet-12M, a dataset containing more than 12 million computed tomography (CT) image slices, to explore the efficacy of self-supervised pretraining on medical and natural images. Our experiments cover intra- and cross-domain transfer scenarios, varying data scales, finetuning vs. linear evaluation, and feature space analysis. We observe that intra-domain transfer compares favorably to cross-domain transfer, achieving comparable or improved performance (0.44% - 2.07% performance increase using RadNet pretraining, depending on the experiment) and demonstrate the existence of a domain boundary-related generalization gap and domain-specific learned features.
Abstract The current gold standard of computer-assisted jaw reconstruction includes raising microvascular bone flaps with patient-specific 3D-printed cutting guides. The downsides of cutting guides are invasive fixation, periosteal denudation, preoperative lead time and missing intraoperative flexibility. This study aimed to investigate the feasibility and accuracy of a robot-assisted cutting method for raising iliac crest flaps compared to a conventional 3D-printed cutting guide. In a randomized crossover design, 40 participants raised flaps on pelvic models using conventional cutting guides and a robot-assisted cutting method. The accuracy was measured and compared regarding osteotomy angle deviation, Hausdorff Distance (HD) and Average Hausdorff Distance (AVD). Duration, workload and usability were further evaluated. The mean angular deviation for the robot-assisted cutting method was 1.9 ± 1.1° (mean ± sd) and for the 3D-printed cutting guide it was 4.7 ± 2.9° (p < 0.001). The HD resulted in a mean value of 1.5 ± 0.6 mm (robot) and 2.0 ± 0.9 mm (conventional) (p < 0.001). For the AVD, this was 0.8 ± 0.5 mm (robot) and 0.8 ± 0.4 mm (conventional) (p = 0.320). Collaborative robot-assisted cutting is an alternative to 3D-printed cutting guides in experimental static settings, achieving slot design benefits with less invasiveness and higher intraoperative flexibility. In the next step, the results should be tested in a dynamic environment with a moving phantom and on the cadaver.