The relative curvature condition (RCC) serves as a crucial constraint, ensuring the avoidance of self-intersection problems in calculating the mean shape over a sample of swept regions. By considering the RCC, this work discusses estimating the mean shape for a class of swept regions called elliptical slabular objects based on a novel shape representation, namely elliptical tube representation (ETRep). The ETRep shape space equipped with extrinsic and intrinsic distances in accordance with object transformation is explained. The intrinsic distance is determined based on the intrinsic skeletal coordinate system of the shape space. Further, calculating the intrinsic mean shape based on the intrinsic distance over a set of ETReps is demonstrated. The proposed intrinsic methodology is applied for the statistical shape analysis to design global and partial hypothesis testing methods to study the hippocampal structure in early Parkinson's disease.
We propose a means of computing fitted frames on the boundary and in the interior of objects and using them to provide the basis for producing geometric features from them that are not only alignment-free but most importantly can be made to correspond locally across a population of objects. We describe a representation targeted for anatomic objects which is designed to enable this strong locational correspondence within object populations and thus to provide powerful object statistics. It accomplishes this by understanding an object as the diffeomorphic deformation of the closure of the interior of an ellipsoid and by using a skeletal representation fitted throughout the deformation to produce a model of the target object, where the object is provided initially in the form of a boundary mesh. Via classification performance on hippocampi shape between individuals with a disorder vs. others, we compare our method to two state-of-theart methods for producing object representations that are intended to capture geometric correspondence across a population of objects and to yield geometric features useful for statistics, and we show notably improved classification performance by this new representation, which we call the evolutionary s-rep. The geometric features that are derived from each of the representations, especially via fitted frames, are discussed.
Monocular depth estimation in endoscopy videos can enable assistive and robotic surgery to obtain better coverage of the organ and detection of various health issues. Despite promising progress on mainstream, natural image depth estimation, techniques perform poorly on endoscopy images due to a lack of strong geometric features and challenging illumination effects. In this paper, we utilize the photometric cues, i.e., the light emitted from an endoscope and reflected by the surface, to improve monocular depth estimation. We first create two novel loss functions with supervised and self-supervised variants that utilize a per-pixel shading representation. We then propose a novel depth refinement network (PPSNet) that leverages the same per-pixel shading representation. Finally, we introduce teacher-student transfer learning to produce better depth maps from both synthetic data with supervision and clinical data with self-supervision. We achieve state-of-the-art results on the C3VD dataset while estimating high-quality depth maps from clinical data. Our code, pre-trained models, and supplementary materials can be found on our project page: https://ppsnet.github.io/
Bronchoscopy is currently the least invasive method for definitively diagnosing lung cancer, which kills more people in the United States than any other form of cancer. Successfully diagnosing suspicious lung nodules requires accurate localization of the bronchoscope relative to a planned biopsy site in the airways. This task is challenging because the lung deforms intraoperatively due to respiratory motion, the airways lack photometric features, and the anatomy's appearance is repetitive. In this paper, we introduce a real-time camera-based method for accurately localizing a bronchoscope with respect to a planned needle insertion pose. Our approach uses deep learning and accounts for deformations and overcomes limitations of global pose estimation by estimating pose relative to anatomical landmarks. Specifically, our learned model considers airway bifurcations along the airway wall as landmarks because they are distinct geometric features that do not vary significantly with respiratory motion. We evaluate our method in a simulated dataset of lungs undergoing respiratory motion. The results show that our method generalizes across patients and localizes the bronchoscope with accuracy sufficient to access the smallest clinically-relevant nodules across all levels of respiratory deformation, even in challenging distal airways. Our method could enable physicians to perform more accurate biopsies and serve as a key building block toward accurate autonomous robotic bronchoscopy.
Statistical analysis of the skeletal structure of slabular objects like groups of hippocampi is valuable for medical researchers as it can be useful for diagnoses and understanding diseases. This work proposes a novel object representation based on model fitting and analysis of the locally parameterized discrete swept skeletal representation of such entities where the model fitting procedure is based on boundary division and surface flattening. The goodness of the model fitting is demonstrated according to the skeletal symmetry, the volume of the implied boundary, and skeletal perturbation. The power of the method is demonstrated by visual inspection and statistical analysis of a synthetic and an actual data set in comparison with an available skeletal representation.
Correlated shape features involving nearby objects often contain important anatomic information. However, it is difficult to capture shape information within and between objects for a joint analysis of multi-object complexes. This paper proposes (1) capturing between-object shape based on an explicit mathematical model called a linking structure, (2) capturing shape features that are invariant to rigid transformation using local affine frames and (3) capturing Correlation of Within- and Between-Object (CoWBO) shape features using a statistical method called NEUJIVE. The resulting correlated shape features give comprehensive understanding of multi-object complexes from various perspectives. First, these features explicitly account for the positional and geometric relations between objects that can be anatomically important. Second, the local affine frames give rise to rich interior geometric features that are invariant to global alignment. Third, the joint analysis of within- and between-object shape yields robust and useful features. To demonstrate the proposed methods, we classify individuals with autism and controls using the extracted shape features of two functionally related brain structures, the hippocampus and the caudate. We found that the CoWBO features give the best classification performance among various choices of shape features. Moreover, the group difference is statistically significant in the feature space formed by the proposed methods.
Three-dimensional (3D) shape lies at the core of understanding the physical objects that surround us. In the biomedical field, shape analysis has been shown to be powerful in quantifying how anatomy changes with time and disease. The Shape AnaLysis Toolbox (SALT) was created as a vehicle for disseminating advanced shape methodology as an open source, free, and comprehensive software tool. We present new developments in our shape analysis software package, including easy-to-interpret statistical methods to better leverage the quantitative information contained in SALT’s shape representations. We also show SlicerPipelines, a module to improve the usability of SALT by facilitating the analysis of large-scale data sets, automating workflows for non-expert users, and allowing the distribution of reproducible workflows.
Reconstructing a 3D surface from colonoscopy video is challenging due to illumination and reflectivity variation in the video frame that can cause defective shape predictions. Aiming to overcome this challenge, we utilize the characteristics of surface normal vectors and develop a two-step neural framework that significantly improves the colonoscopy reconstruction quality. The normal-based depth initialization network trained with self-supervised normal consistency loss provides depth map initialization to the normal-depth refinement module, which utilizes the relationship between illumination and surface normals to refine the frame-wise normal and depth predictions recursively. Our framework’s depth accuracy performance on phantom colonoscopy data demonstrates the value of exploiting the surface normals in colonoscopy reconstruction, especially on en face views. Due to its low depth error, the prediction result from our framework will require limited post-processing to be clinically applicable for real-time colonoscopy reconstruction.
Objects and object complexes in 3D, as well as those in 2D, have many possible representations. Among them skeletal representations have special advantages and some limitations. For the special form of skeletal representation called "s-reps," these advantages include strong suitability for representing slabular object populations and statistical applications on these populations. Accomplishing these statistical applications is best if one recognizes that s-reps live on a curved shape space. Here we will lay out the definition of s-reps, their advantages and limitations, their mathematical properties, methods for fitting s-reps to single- and multi-object boundaries, methods for measuring the statistics of these object and multi-object representations, and examples of such applications involving statistics. While the basic theory, ideas, and programs for the methods are described in this paper and while many applications with evaluations have been produced, there remain many interesting open opportunities for research on comparisons to other shape representations, new areas of application and further methodological developments, many of which are explicitly discussed here.
Shape correlation of multi-object complexes in the human body can have significant implications in understanding the development of disease. While there exist geometric and statistical methods that aim for multi-object shape analysis, very little research can effectively extract shape correlation. It is especially difficult to extract the correlation when the involved objects have different variability in separate non-Euclidean spaces. To address these difficulties, this paper proposes geometric and statistical methods to extract the shape correlation from multi-object complexes. In particular, we focus on the shape correlation of the hippocampus and the caudate subject to the development of autism. The proposed methods are designed (1) to capture objects’ shape features (2) to capture shape correlation regardless of different variability between the two objects and (3) to provide interpretable shape correlation in multi-object complexes. In our experiments on synthetic data and autism data, the quantitative results and the qualitative visualization suggest that our methods are effective and robust.
One of the key elements of reconstructing a 3D mesh from a monocular video is generating every frame's depth map. However, in the application of colonoscopy video reconstruction, producing good-quality depth estimation is challenging. Neural networks can be easily fooled by photometric distractions or fail to capture the complex shape of the colon surface, predicting defective shapes that result in broken meshes. Aiming to fundamentally improve the depth estimation quality for colonoscopy 3D reconstruction, in this work we have designed a set of training losses to deal with the special challenges of colonoscopy data. For better training, a set of geometric consistency objectives was developed, using both depth and surface normal information. Also, the classic photometric loss was extended with feature matching to compensate for illumination noise. With the training losses powerful enough, our self-supervised framework named ColDE is able to produce better depth maps of colonoscopy data as compared to the previous work utilizing prior depth knowledge. Used in reconstruction, our network is able to reconstruct good-quality colon meshes in real-time without any post-processing, making it the first to be clinically applicable.
Place recognition in colonoscopy is needed for various reasons. 1) If a certain region needs to be rechecked during an endoscopy, the endoscopist needs to re-localize the camera accurately to the region of interest. 2) Place recognition is needed for same-patient follow-up colonoscopy to localize the region where a polyp was cut off. 3) Recent development in colonoscopic 3D reconstruction needs place recognition to establish long-range correspondence, e.g., for loop closure. However, traditional image retrieval techniques do not generalize well in colonic images. Moreover, although place recognition or instance-level image retrieval is a widely researched topic in computer vision and several benchmarks have been published for it, there has been no specific research or benchmarks in endoscopic images, which are significantly different from common images used in traditional computer vision tasks. In this paper we present a testing dataset with manually labeled groundtruth which comprises 10126 images from 20 colonoscopic subsequences. We perform an extensive evaluation on different existing place recognition techniques using different metrics.
High screening coverage during colonoscopy is crucial to effectively prevent colon cancer. Previous work has allowed alerting the doctor to unsurveyed regions by reconstructing the 3D colonoscopic surface from colonoscopy videos in real-time. However, the lighting inconsistency of colonoscopy videos can cause a key component of the colonoscopic reconstruction system, the SLAM optimization, to fail. In this work we focus on the lighting problem in colonoscopy videos. To successfully improve the lighting consistency of colonoscopy videos, we have found necessary a lighting correction that adapts to the intensity distribution of recent video frames. To achieve this in real-time, we have designed and trained an RNN network. This network adapts the gamma value in a gamma-correction process. Applied in the colonoscopic surface reconstruction system, our light-weight model significantly boosts the reconstruction success rate, making a larger proportion of colonoscopy video segments reconstructable and improving the reconstruction quality of the already reconstructed segments.
Colonoscopy is the gold standard for pre-cancerous polyps screening and treatment. The polyp detection rate is highly tied to the percentage of surveyed colonic surface. However, current colonoscopy technique cannot guarantee that all the colonic surface is well examined because of incomplete camera orientations and of occlusions. The missing regions can hardly be noticed in a continuous first-person perspective. Therefore, a useful contribution would be an automatic system that can compute missing regions from an endoscopic video in real-time and alert the endoscopists when a large missing region is detected. We present a novel method that reconstructs dense chunks of a 3D colon in real time, leaving the unsurveyed part unreconstructed. The method combines a standard SLAM system with a depth and pose prediction network to achieve much more robust tracking and less drift. It addresses the difficulties for colonoscopic images of existing simultaneous localization and mapping (SLAM) systems and end-to-end deep learning methods.
This paper considers joint analysis of multiple functionally related structures in classification tasks. In particular, our method developed is driven by how functionally correlated brain structures vary together between autism and control groups. To do so, we devised a method based on a novel combination of (1) non-Euclidean statistics that can faithfully represent non-Euclidean data in Euclidean spaces and (2) a non-parametric integrative analysis method that can decompose multi-block Euclidean data into joint, individual, and residual structures. We find that the resulting joint structure is effective, robust, and interpretable in recognizing the underlying patterns of the joint variation of multi-block non-Euclidean data. We verified the method in classifying the structural shape data collected from cases that developed and did not develop into Autistic Spectrum Disorder (ASD).
Representing an object by a skeletal structure can be powerful for statistical shape analysis if there is good correspondence of the representations within a population. Many anatomic objects have a genus-zero boundary and can be represented by a smooth unbranching skeletal structure that can be discretely approximated. We describe how to compute such a discrete skeletal structure ("d-s-rep") for an individual 3D shape with the desired correspondence across cases. The method involves fitting a d-s-rep to an input representation of an object's boundary. A good fit is taken to be one whose skeletally implied boundary well approximates the target surface in terms of low order geometric boundary properties: (1) positions, (2) tangent fields, (3) various curvatures. Our method involves a two-stage framework that first, roughly yet consistently fits a skeletal structure to each object and second, refines the skeletal structure such that the shape of the implied boundary well approximates that of the object. The first stage uses a stratified diffeomorphism to produce topologically non-self-overlapping, smooth and unbranching skeletal structures for each object of a population. The second stage uses loss terms that measure geometric disagreement between the skeletally implied boundary and the target boundary and avoid self-overlaps in the boundary. By minimizing the total loss, we end up with a good d-s-rep for each individual shape. We demonstrate such d-s-reps for various human brain structures. The framework is accessible and extensible by clinical users, researchers and developers as an extension of SlicerSALT, which is based on 3D Slicer.
Cerebrospinal fluid (CSF) plays an essential role in early postnatal brain development. Extra-axial CSF (EA-CSF) volume, which is characterized by CSF in the subarachnoid space surrounding the brain, is a promising marker in the early detection of young children at risk for neurodevelopmental disorders. Previous studies have focused on global EA-CSF volume across the entire dorsal extent of the brain, and not regionally-specific EA-CSF measurements, because no tools were previously available for extracting local EA-CSF measures suitable for localized cortical surface analysis. In this paper, we propose a novel framework for the localized, cortical surface-based analysis of EA-CSF. The proposed processing framework combines probabilistic brain tissue segmentation, cortical surface reconstruction, and streamline-based local EA-CSF quantification. The quantitative analysis of local EA-CSF was applied to a dataset of typically developing infants with longitudinal MRI scans from 6 to 24 months of age. There was a high degree of consistency in the spatial patterns of local EA-CSF across age using the proposed methods. Statistical analysis of local EA-CSF revealed several novel findings: several regions of the cerebral cortex showed reductions in EA-CSF from 6 to 24 months of age, and specific regions showed higher local EA-CSF in males compared to females. These age-, sex-, and anatomically-specific patterns of local EA-CSF would not have been observed if only a global EA-CSF measure were utilized. The proposed methods are integrated into a freely available, open-source, cross-platform, user-friendly software tool, allowing neuroimaging labs to quantify local extra-axial CSF in their neuroimaging studies to investigate its role in typical and atypical brain development.
Deep learning-based, single-view depth estimation methods have recently shown highly promising results. However, such methods ignore one of the most important features for determining depth in the human vision system, which is motion. We propose a learning-based, multi-view dense depth map and odometry estimation method that uses Recurrent Neural Networks (RNN) and trains utilizing multi-view image reprojection and forward-backward flow-consistency losses. Our model can be trained in a supervised or even unsupervised mode. It is designed for depth and visual odometry estimation from video where the input frames are temporally correlated. However, it also generalizes to single-view depth estimation. Our method produces superior results to the state-of-the-art approaches for single-view and multi-view learning-based depth estimation on the KITTI driving dataset.
Skeletal models that are structurally medial provide effective object representations. This is because they include not only locations but also boundary directions and object widths. We present a skeletal object representation that we call "quasimedial" because geometric properties associated with Blum's [1] medial axis are relaxed to allow the skeleton to have a pre-specified amount of branching and thus to support statistical analysis. We call this form of object representation the s-rep. We explain how such models can be automatically determined from object boundary data in a way that a) avoids boundary noise, b) implies a boundary that closely fits the input boundary, and c) well recognizes shape correspondences across cases. We also explain how to use Riemannian geometry to estimate probability distributions from a sample of s-reps and to find ways to classify s-reps between two categories as trained from s-reps in each class. Finally, we describe various evidence that shows the relative strengths of s-reps vs. other object representations; we also discuss shortcomings of s-reps.