The neuropsychological crowding effect denotes the reallocation of cognitive functions within the contralesional hemisphere following unilateral brain damage, prioritizing language at the expense of nonverbal abilities. This study investigates structural white matter correlates of crowding in the arcuate fasciculus (AF), a key language tract, using hemispherotomy as a unique setting to explore structural reorganization supporting language preservation. We explore two main hypotheses. First, the contralesional right AF undergoes white matter reorganization correlated with preserved language function at the expense of nonverbal abilities following left-hemispheric damage. Second, this reorganization varies with epilepsy etiology, influencing different stages of developmental language lateralization. This retrospective study included individuals post-hemispherotomy and healthy controls. Inclusion criteria were; (1) being a native German speaker, (2) having no MRI contraindication, (3) the ability to undergo approximately 2 h of MRI scans, and (4) the ability to participate in neuropsychological assessments over two consecutive days. Neuroimaging included T1-, T2-, and diffusion-weighted imaging, alongside postoperative neuropsychological assessments, where it was taken as evidence for crowding if verbal IQ exceeded performance IQ by at least 10 points. The AF was reconstructed using advanced tractography, and CoBundleMAP was used to compare morphologically corresponding AF subsections. Statistical significance was set at p < 0.05 $$ p<0.05 $$ , with correction for multiple comparisons applied across contiguous tract sections using Threshold-Free Cluster Enhancement. The final cohort comprised 22 individuals post-hemispherotomy (median age: 20.4 $$ 20.4 $$ years, range: 12.3 - 43.9 $$ 12.3-43.9 $$ ; 55% female; 55% with left-sided surgeries) and 20 healthy controls (median age: 23.8 $$ 23.8 $$ years, range: 15.5 - 54.0 $$ 15.5-54.0 $$ ; 55% female). Crowding was associated with significantly higher fractional anisotropy (FA) in the AF ( p = 0.015 $$ p=0.015 $$ , Cohen's d = 1.69 $$ d=1.69 $$ ), but only observed in individuals with left-sided hemispherotomy, localized to a subsection between Geschwind's territory and Wernicke's area ( p corrected = 0.02 $$ {p}_{\mathrm{corrected}}=0.02 $$ ). This region also displayed significantly higher normalized FA in AF of individuals with congenital etiology and crowding compared to acquired etiology and no crowding ( p corrected = 0.0189 $$ {p}_{\mathrm{corrected}}=0.0189 $$ ). This study identifies previously unreported neural correlates of crowding in right contralesional AF of individuals post-hemispherotomy and highlights specific AF subsections involved in preserving language functions at the cost of nonverbal abilities. The findings suggest a link between crowding and epilepsy etiology, particularly in the region spanning Geschwind's territory and Wernicke's area.
Weakly supervised segmentation has the potential to greatly reduce the annotation effort for training segmentation models for small structures such as hyper-reflective foci (HRF) in optical coherence tomography (OCT). However, most weakly supervised methods either involve a strong downsampling of input images, or only achieve localization at a coarse resolution, both of which are unsatisfactory for small structures. We propose a novel framework that increases the spatial resolution of a traditional attention-based Multiple Instance Learning (MIL) approach by using Layer-wise Relevance Propagation (LRP) to prompt the Segment Anything Model (SAM 2), and increases recall with iterative inference. Moreover, we demonstrate that replacing MIL with a Compact Convolutional Transformer (CCT), which adds a positional encoding, and permits an exchange of information between different regions of the OCT image, leads to a further and substantial increase in segmentation accuracy.
The facial gestalt (overall facial morphology) is a characteristic clinical feature in many genetic disorders that is often essential for suspecting and establishing a specific diagnosis. Therefore, publishing images of individuals affected by pathogenic variants in disease-associated genes has been an important part of scientific communication. Furthermore, medical imaging data is also crucial for teaching and training deep-learning models such as GestaltMatcher. However, medical data is often sparsely available, and sharing patient images involves risks related to privacy and re-identification. Therefore, we explored whether generative neural networks can be used to synthesize accurate portraits for rare disorders. We modified a StyleGAN architecture and trained it to produce artificial condition-specific portraits for multiple disorders. In addition, we present a technique that generates a sharp and detailed average patient portrait for a given disorder. We trained our GestaltGAN on the 20 most frequent disorders from the GestaltMatcher database. We used REAL-ESRGAN to increase the resolution of portraits from the training data with low-quality and colorized black-and-white images. To augment the model’s understanding of human facial features, an unaffected class was introduced to the training data. We tested the validity of our generated portraits with 63 human experts. Our findings demonstrate the model’s proficiency in generating photorealistic portraits that capture the characteristic features of a disorder while preserving patient privacy. Overall, the output from our approach holds promise for various applications, including visualizations for publications and educational materials and augmenting training data for deep learning.
Whole-brain tractography in diffusion MRI is often followed by a parcellation in which each streamline is classified as belonging to a specific white matter bundle, or discarded as a false positive. Efficient parcellation is important both in large-scale studies, which have to process huge amounts of data, and in the clinic, where computational resources are often limited. TractCloud is a state-of-the-art approach that aims to maximize accuracy with a local-global representation. We demonstrate that the local context does not contribute to the accuracy of that approach, and is even detrimental when dealing with pathological cases. Based on this observation, we propose PETParc, a new method for Parallel Efficient Tractography Parcellation. PETParc is a transformer-based architecture in which the whole-brain tractogram is randomly partitioned into sub-tractograms whose streamlines are classified in parallel, while serving as global context for each other. This leads to a speedup of up to two orders of magnitude relative to TractCloud, and permits inference even on clinical workstations without a GPU. PETParc accounts for the lack of streamline orientation either via a novel flip-invariant embedding, or by simply using flips as part of data augmentation. Despite the speedup, results are often even better than those of prior methods. The code and pretrained model will be made public upon acceptance.
Deep learning-based tractography implicitly learns anatomical prior knowledge that is required to resolve ambiguities inherent in traditional streamline tractography. TractSeg is a particularly widely used example of such an approach. Even though it has exclusively been trained on healthy subjects, a certain level of generalization to different pathologies has been demonstrated, and TractSeg is now increasingly used for clinical cases. We explore the limits of TractSeg by evaluating it on a unique dataset of 25 patients with epilepsy who underwent hemispherotomy, a type of surgery in which the two hemispheres are surgically separated. We compare results to those on 25 healthy controls who have been imaged with the same setup.We find that TractSeg generalizes remarkably well, given the severity of the abnormalities. However, to our knowledge, we are the first to document cases in which TractSeg erroneously reconstructs (“hallucinates”) tracts that are known to have been surgically disconnected, and we found cases in which it implausibly continues tracts through obvious lesions. At the same time, TractSeg failed to reconstruct or undersegmented some tracts that are known to be preserved.We subsequently propose a refinement of TractSeg which aims to improve its applicability to data with pathologies, by using its Tract Orientation Maps as an anatomical prior in low-rank tensor approximation based tractography such that tracking is guaranteed to continue only where presence of the tract is directly supported by the data (“data fidelity”). We demonstrate that our extension not only eliminates hallucinated tracts and reconstructions within lesions, but that it also increases the ability to reconstruct the preserved tracts, and leads to more complete reconstructions even in healthy controls. Despite these advances, we recommend caution and manual quality control when applying deep learning based tractography to patient data.
Manual Small-Incision Cataract Surgery (SICS) is a prevalent technique in low- and middle-income countries (LMICs) but understudied with respect to computer assisted surgery. This prospective cross-sectional study introduces the first SICS video dataset, evaluates effectiveness of phase recognition through deep learning (DL) using the MS-TCN + + architecture, and compares its results with the well-studied phacoemulsification procedure using the Cataract-101 public dataset. Our novel SICS-105 dataset involved 105 patients recruited at Sankara Eye Hospital in India. Performance is evaluated with frame-wise accuracy, edit distance, F1-score, Precision-Recall AUC, sensitivity, and specificity. The MS-TCN + + architecture performs better on the Cataract-101 dataset, with an accuracy of 89.97% [CI 86.69–93.46%] compared to 85.56% [80.63–92.09%] on the SICS-105 dataset (ROC AUC 99.10% [98.34–99.51%] vs. 98.22% [97.16–99.26%]). The accuracy distribution and confidence-intervals overlap and the ROC AUC values range 46.20 to 94.18%. Even though DL is found to be effective for phase recognition in SICS, the larger number of phases and longer duration makes it more challenging compared to phacoemulsification. To support further developments, we make our dataset open access. This research marks a crucial step towards improving postoperative analysis and training for SICS.
Cataract surgery is the most common surgical procedure globally, with a disproportionately higher burden in developing countries. While automated surgical video analysis has been explored in general surgery, its application to ophthalmic procedures remains limited. Existing research primarily focuses on Phaco cataract surgery, an expensive technique not accessible in regions where cataract treatment is most needed. In contrast, Manual Small-Incision Cataract Surgery (MSICS) is the preferred low-cost alternative in high-volume settings and for complex cases. However, no dataset exists for MSICS. To address this gap, we introduce Sankara-MSICS, the first comprehensive dataset containing 53 surgical videos annotated for 18 surgical phases and 3,527 frames with 13 surgical tools at the pixel level. We also present ToolSeg, a novel framework that enhances tool segmentation with a phase-conditional decoder and a semi-supervised setup leveraging pseudo-labels from foundation models. Our approach significantly improves segmentation performance, achieving a 38.1% increase in mean Dice scores, with notable gains for smaller and less prevalent tools. The code is available at https://github.com/Sri-Kanchi-Kamakoti-Medical-Trust/ToolSeg.
The foveola, the central region of the human retina, plays a crucial role in sharp color vision and is challenging to study due to its unique anatomy and technical limitations in imaging. We present ConeMapper, an open-source MATLAB software that integrates a fully convolutional neural network (FCN) for the automatic detection and analysis of cone photoreceptors in confocal adaptive optics scanning light ophthalmoscopy (AOSLO) images of the foveal center. The FCN was trained on a dataset of 49 healthy retinas and showed improved performance over previously published neural networks, particularly in the central fovea, achieving an $$F_1$$ score of 0.9769 across the validation set, critically reducing analysis time. In addition to automatic cone detection, ConeMapper provides efficient manual annotation tools, visualizations and topographical analysis, offering users detailed metrics for further analysis. ConeMapper is freely available, with ongoing development aimed at enhancing functionality and adaptability to different retinal imaging modalities.
Fast algorithms for diffusion MRI tractography are required due to the increasing amounts of diffusion MRI data, and the increasing popularity of whole-brain tractography. Representing fiber orientation density functions (fODFs) as higher-order tensors and extracting main fiber directions from them via low-rank tensor approximation is a state-of-the-art variant of streamline tractography, but involves a computationally costly nonlinear optimization in each integration step. In this work, we demonstrate that unsupervised training of a neural network to map fODF coefficients to the corresponding fiber contributions directly is not only faster, but also achieves lower approximation residuals. This is due to the fact that training the network amounts to a joint optimization of all fiber contributions, while traditional algorithms follow an alternating optimization strategy. However, we observe that the traditional approach implicitly favors sparse solutions, and that a corresponding explicit regularization is required to obtain useful results with a joint optimization strategy. Building on those insights, we create the first GPU-based implementation of low-rank tractography, which achieves a speedup by a factor of 68, compared to traditional tractography on a single CPU core, while at the same time improving the median dice.
Low-rank higher-order tensor approximation has been used successfully to extract discrete directions for tractography from continuous fiber orientation density functions (fODFs). However, while it accounts for fiber crossings, it has so far ignored fanning, which has led to incomplete reconstructions. In this work, we integrate an anisotropic model of fanning based on the Bingham distribution into a recently proposed tractography method that performs low-rank approximation with an Unscented Kalman Filter. Our technical contributions include an initialization scheme for the new parameters, which is based on the Hessian of the low-rank approximation, pre-integration of the required convolution integrals to reduce the computational effort, and representation of the required 3D rotations with quaternions. Results on 12 subjects from the Human Connectome Project confirm that, in almost all considered tracts, our extended model significantly increases completeness of the reconstruction, at acceptable excess and additional computational cost. Its results are also more accurate than those from a simpler, isotropic fanning model that is based on Watson distributions.
We propose an algorithmic pipeline that uses interpretable Generative Adversarial Networks (GANs) to visualize the variability of the visual appearance of drusen in Optical Coherence Tomography (OCT). Drusen are accumulations of extracellular debris between Bruch's membrane and the retinal pigment epithelium of the eye. They are a hallmark of age-related macular degeneration (AMD)-the most common cause of vision loss in the elderly. Imaging the morphology of drusen with OCT reveals different subtypes, which might have different relevance for disease severity and the risk of progression. We compare two GAN architectures and three recently proposed methods for the unsupervised discovery of interpretable paths in their latent space with respect to their ability to visualize natural variations in drusen appearance. We also introduce a color code that indicates generated images that extrapolate beyond the training data and should, therefore, be interpreted with caution. Our results suggest that, even when trained on cross-sectional data, GANs can recover smooth and anatomically plausible variations of drusen that are in agreement with changes over time that are known from longitudinal observations.
Scribbles are a popular form of weak annotation for the segmentation of three-dimensional medical images, but typically require iterative refinement to achieve the desired segmentation map. The complexity of diffusion MRI (dMRI) poses additional challenges. Previous work addressed the high dimensionality of dMRI via unsupervised representation learning, and combined it with a random forest classifier that can be re-trained quickly enough to provide interactive feedback to the human annotator. Our work extends that framework in multiple ways. Our main contribution is to add an active learning component that suggests locations in which additional scribbles should be placed. It relies on uncertainty quantification via test time augmentation (TTA). Second, we observe that TTA increases segmentation accuracy even by itself. Moreover, we demonstrate that anomaly detection via isolation forests effectively suppresses false positives that arise when generalizing from sparse scribbles. Taken together, these contributions substantially improve the accuracy that can be achieved with various annotation budgets.
Purpose:The purpose of this study was to assess the current use and reliability of artificial intelligence (AI)-based algorithms for analyzing cataract surgery videos. Methods:A systematic review of the literature about intra-operative analysis of cataract surgery videos with machine learning techniques was performed. Cataract diagnosis and detection algorithms were excluded. Resulting algorithms were compared, descriptively analyzed, and metrics summarized or visually reported. The reproducibility and reliability of the methods and results were assessed using a modified version of the Medical Image Computing and Computer-Assisted (MICCAI) checklist. Results:Thirty-eight of the 550 screened studies were included, 20 addressed the challenge of instrument detection or tracking, 9 focused on phase discrimination, and 8 predicted skill and complications. Instrument detection achieves an area under the receiver operator characteristic curve (ROC AUC) between 0.976 and 0.998, instrument tracking an mAP between 0.685 and 0.929, phase recognition an ROC AUC between 0.773 and 0.990, and complications or surgical skill performs with an ROC AUC between 0.570 and 0.970. Conclusions:The studies showed a wide variation in quality and pose a challenge regarding replication due to a small number of public datasets (none for manual small incision cataract surgery) and seldom published source code. There is no standard for reported outcome metrics and validation of the models on external datasets is rare making comparisons difficult. The data suggests that tracking of instruments and phase detection work well but surgical skill and complication recognition remains a challenge for deep learning. Translational Relevance:This overview of cataract surgery analysis with AI models provides translational value for improving training of the clinician by identifying successes and challenges.
Domain shift occurs when training U-Nets for medical image segmentation with images from one device, but applying them to images from a different device. This often reduces accuracy, and it poses a challenge for uncertainty quantification, when incorrect segmentations are produced with high confidence. Recent work proposed to detect such failure cases via anomalies in feature space: Activation patterns that deviate from those observed during training are taken as an indication that the input is not handled well by the network, and its output should not be trusted. However, such latent space distances primarily detect whether images are from different scanners, not whether they are correctly segmented. Therefore, we propose a novel segmentation distortion measure for uncertainty quantification. It uses an autoencoder to make activations more similar to those that were observed during training, and propagates the result through the remainder of the U-Net. We demonstrate that the extent to which this affects the segmentation correlates much more strongly with segmentation errors than distances in activation space, and that it quantifies uncertainty under domain shift better than entropy in the output of a single U-Net, or an ensemble of U-Nets.
In clinical practice, Diffusion Tensor Magnetic Resonance Imaging (DT-MRI) is usually evaluated by visual inspection of grayscale maps of Fractional Anisotropy or mean diffusivity. However, the fact that those maps only contain part of the information that is captured in DT-MRI implies a risk of missing signs of disease. In this work, we propose a visualization system that supports a more comprehensive analysis with an anomaly score that accounts for the full diffusion tensor information. It is computed by comparing the DT-MRI scan of a given patient to a control group of healthy subjects, after spatial coregistration. Moreover, our system introduces an Anomaly Lens which visualizes how a user-specified region of interest deviates from the controls, indicating which aspects of the tensor (norm, anisotropy, mode, rotation) differ most, whether they are elevated or reduced, and whether their covariation matches the covariances within the control group. Applying our system to patients with metachromatic leukodystrophy clearly indicates regions affected by the disease, and permits their detailed analysis.
Diffusion MRI is a modern neuroimaging modality with a unique ability to acquire microstructural information by measuring water self-diffusion at the voxel level. However, it generates huge amounts of data, resulting from a large number of repeated 3D scans. Each volume samples a location in q-space, indicating the direction and strength of a diffusion sensitizing gradient during the measurement. This captures detailed information about the self-diffusion, and the tissue microstructure that restricts it. Lossless compression with GZIP is widely used to reduce the memory requirements. We introduce a novel lossless codec for diffusion MRI data. It reduces file sizes by more than 30% compared to GZIP, and also beats lossless codecs from the JPEG family. Our codec builds on recent work on lossless PDE-based compression of 3D medical images, but additionally exploits smoothness in q-space. We demonstrate that, compared to using only image space PDEs, q-space PDEs further improve compression rates. Moreover, implementing them with Finite Element Methods and a custom acceleration significantly reduces computational expense. Finally, we show that our codec clearly benefits from integrating subject motion correction, and slightly from optimizing the order in which the 3D volumes are coded.
Tractography based on diffusion Magnetic Resonance Imaging (dMRI) is the prevalent approach to the in vivo delineation of white matter tracts in the human brain. Many tractography methods rely on models of multiple fiber compartments, but the local dMRI information is not always sufficient to reliably estimate the directions of secondary fibers. Therefore, we introduce two novel approaches that use spatial regularization to make multi-fiber tractography more stable. Both represent the fiber Orientation Distribution Function (fODF) as a symmetric fourth-order tensor, and recover multiple fiber orientations via low-rank approximation. Our first approach computes a joint approximation over suitably weighted local neighborhoods with an efficient alternating optimization. The second approach integrates the low-rank approximation into a current state-of-the-art tractography algorithm based on the unscented Kalman filter (UKF). These methods were applied in three different scenarios. First, we demonstrate that they improve tractography even in high-quality data from the Human Connectome Project, and that they maintain useful results with a small fraction of the measurements. Second, on the 2015 ISMRM tractography challenge, they increase overlap, while reducing overreach, compared to low-rank approximation without joint optimization or the traditional UKF, respectively. Finally, our methods permit a more comprehensive reconstruction of tracts surrounding a tumor in a clinical dataset. Overall, both approaches improve reconstruction quality. At the same time, our modified UKF significantly reduces the computational effort compared to its traditional counterpart, and to our joint approximation. However, when used with ROI-based seeding, joint approximation more fully recovers fiber spread.
Diffusion magnetic resonance imaging (dMRI) tractography has the unique ability to reconstruct major white matter tracts non‐invasively and is, therefore, widely used in neurosurgical planning and neuroscience. In this work, we reduce two sources of uncertainty within the tractography pipeline. The first one is the model uncertainty that arises in crossing fibre tractography, from having to estimate the number of relevant fibre compartments in each voxel. We propose a mathematical framework to estimate model uncertainty, and we reduce this type of uncertainty with a model averaging approach that combines the fibre direction estimates from all candidate models, weighted by the posterior probability of the respective model. The second source of uncertainty is measurement noise. We use bootstrapping to estimate this data uncertainty, and consolidate the fibre direction estimates from all bootstraps into a consensus model. We observe that, in most voxels, a traditional model selection strategy selects different models across bootstraps. In this sense, the bootstrap consensus also reduces model uncertainty. Either approach significantly increases the accuracy of crossing fibre tractography in multiple subjects, and combining them provides an additional benefit. However, model averaging is much more efficient computationally.
Drusen are an important biomarker for age-related macular degeneration (AMD). Their accurate segmentation based on optical coherence tomography (OCT) is therefore relevant to the detection, staging, and treatment of disease. Since manual OCT segmentation is resource-consuming and has low reproducibility, automatic techniques are required. In this work, we introduce a novel deep learning based architecture that directly predicts the position of layers in OCT and guarantees their correct order, achieving state-of-the-art results for retinal layer segmentation. In particular, the average absolute distance between our model’s prediction and the ground truth layer segmentation in an AMD dataset is 0.63, 0.85, and 0.44 pixel for Bruch's membrane (BM), retinal pigment epithelium (RPE) and ellipsoid zone (EZ), respectively. Based on layer positions, we further quantify drusen load with excellent accuracy, achieving 0.994 and 0.988 Pearson correlation between drusen volumes estimated by our method and two human readers, and increasing the Dice score to 0.71 ± 0.16 (from 0.60 ± 0.23) and 0.62 ± 0.23 (from 0.53 ± 0.25), respectively, compared to a previous state-of-the-art method. Given its reproducible, accurate, and scalable results, our method can be used for the large-scale analysis of OCT data.
Background: Smartphone-based fundus imaging (SBFI) is an innovative and low-cost alternative for color fundus photography. Since the first reports on this topic more than 10 years ago a large number of studies on different adapters and clinical applications have been published. Objective: The aim of this review article is to provide an overview on the development of SBFI and adapters and clinical applications published so far. Material and methods: A literature search was performed using the MEDLINE and Science Citation Index Expanded databases without time restrictions. Results: Overall, 11 adapters were included and compared in terms of exemplary image material, field of view, acquisition costs, weight, software, application range, smartphone compatibility and certification. Previously published SBFI applications are screening for diabetic retinopathy, glaucoma and retinopathy of prematurity as well as the application in emergency medicine, pediatrics and medical education/teaching. Image quality of conventional retinal cameras is in general superior to SBFI. First approaches on automatic detection of diabetic retinopathy through SBFI are promising and the use of automatic image processing algorithms enables the generation of widefield image montages. Conclusion: SBFI is a versatile, mobile, low-cost alternative to conventional equipment for color fundus photography. In addition, it facilitates the delegation of ophthalmological examinations to assistance personnel in telemedical settings, could simplify retinal documentation, improve teaching, and improve ophthalmological care, particularly in countries with low and middle incomes.