Objective: Automated medical image segmentation (MIS) using deep learning has traditionally relied on models built and trained from scratch, or at least fine-tuned on a target dataset. The Segment Anything Model (SAM) by Meta challenges this paradigm by providing zero-shot generalisation capabilities. This study aims to develop and compare methods for refining traditional U-Net segmentations by repurposing them for automated SAM prompting. Approach: A 2D U-Net with EfficientNet-B4 encoder was trained using 4-fold cross-validation on an in-house brain metastases dataset. Segmentation predictions from each validation set were used for automatic sparse prompt generation via a bounding box prompting method (BBPM) and novel implementations of the point prompting method (PPM). The PPMs frequently produced poor slice predictions (PSPs) that required identification and substitution. A slice was identified as a PSP if it (1) contained multiple predicted regions per lesion or (2) possessed outlier foreground pixel counts relative to the patient's other slices. Each PSP was substituted with a corresponding initial U-Net or SAM BBPM prediction. The patients' mean volumetric dice similarity coefficient (DSC) was used to evaluate and compare the methods' performances. Main results: Relative to the initial U-Net segmentations, the BBPM improved mean patient DSC by 3.93 +/- 1.48% to 0.847 +/- 0.008 DSC. PSPs constituted 20.01-21.63% of PPMs' predictions and without substitution performance dropped by 82.94 +/- 3.17% to 0.139 +/- 0.023 DSC. Pairing the two PSP identification techniques yielded a sensitivity to PSPs of 92.95 +/- 1.20%. By combining this approach with BBPM prediction substitution, the PPMs achieved segmentation accuracies on par with the BBPM, improving mean patient DSC by up to 4.17 +/- 1.40% and reaching 0.849 +/- 0.007 DSC. Significance: The proposed PSP identification and substitution techniques bridge the gap between PPM and BBPM performance for MIS. Additionally, the uniformity observed in our experiments' results demonstrates the robustness of SAM to variations in prompting style. These findings can assist in the design of both automatically and manually prompted pipelines.
Background/Aim: The problem of lack of standardisation in target delineation and herewith the variability of target contours in Gamma Knife radiosurgery is as severe as in linac-based radiotherapy in general. The first aim of this study was to quantify the contouring variability for a group of five radiosurgery targets and estimate their true-volume based on multiple delineations using the Simultaneous Truth and Performance Level Estimation (STAPLE) algorithm. The second aim was to assess the robustness of the STAPLE method for the assessment of the true-volume, with respect to the number of contours available as input. Patients and Methods: A multicentre analysis of the variability in contouring of five cases was performed. Twelve contours were provided for each case by experienced planners for Gamma Knife. To assess the robustness of the STAPLE method with respect to the number of contours used as input, sets of contours were randomly selected in the analysis. Results: A high similarity was observed between the STAPLE generated true-volume and the 50%-agreement volume when all 12 available contours were used as input (90-100%). Lower similarity was observed with smaller sets of contours (10-70%). Conclusion: If a high number of input contours is available, the STAPLE method provides a valuable tool in the estimation of the true volume of a target based on multiple contours as well as the sensitivity and specificity for each input contour relative to the true volume of that structure. The robustness of the STAPLE method for rendering the true target volume depends on the number of contours provided as input and their variability with respect to shape, size and position.
Radiosurgery (RS) treatment times vary, even for the same prescription dose, due to variations in the collimator size, the number of iso-centres/beams/arcs used and the time gap between each of these exposures. The biologically effective dose (BED) concept, incorporating fast and slow components of repair, was used to show the likely influence of these variables for Gamma Knife patients with Vestibular Schwannomas. Two patients plans were selected, treated with the Model B Gamma Knife, these representing the widest range of treatment variables; iso-centre numbers 3 and 13, overall treatment times 25.4 and 129.6 min, prescription dose 14 Gy. These were compared with 3 cases treated with the Perfexion (R) Gamma Knife. The iso-centre number varied between 11 and 18, treatment time 35.7 - 74.4 min, prescription dose 13 Gy. In the longer Model B Gamma Knife treatment plan the 14 Gy iso-dose was best matched by the 58 Gy(2.47) iso-BED line, although higher and lower BED values were associated with regions on the prescription iso-dose. The equivalent value for the shorter treatment was 85 Gy(2.47). BED volume histograms showed that a BED of 85 Gy(2.47) only covered similar to 65% of the target in the plan with the longer overall treatment time. The corresponding BED values for the 3 cases, treated with the Perfexion (R) Gamma Knife, were 59.5, 68.5 and 71.5 Gy(2.47).In conclusion BED calculations, taking account of the repair of sublethal damage, may indicate the importance of reporting overall time to reflect the biological effectiveness of the total physical dose applied. (C) 2015 Published by Elsevier Ltd on behalf of Associazione Italiana di Fisica Medica.
In the application of stereotactic radiosurgery, using the Gamma Knife, there are large variations in the overall treatment time for the same prescription dose, given in a single treatment session, for different patients. This is due to not only changes in the activity of the Cobolt-60 sources, but also to variations in the number of iso-centers used, the collimator size for a particular iso-center, and the time gap between the different iso-centers. Although frequently viewed as a single dose treatment the concept of biologically effective dose (BED), incorporating concurrent fast and a slow components of repair of sublethal damage, would imply potential variations in BED because of the influence of these different variables associated with treatment. This was investigated in 26 patients, treated for Vestibular Schwannomas, using the Series B Gamma-Knife, between 1999 and 2005. The iso-center number varied between 2 and 13, and the overall treatment time from 25.4-129.58 min. The prescription doses varied from 10-14 Gy. To obtain physical dose and dose-rates from each iso-center, in a number of locations in the region of interest, a prototype version of the Leksell GammaPlan((R)) was used. For an individual patient, BED values varied by up to 15% for a given physical iso-dose. This was due to variation in the dose prescription at different locations on that iso-dose. Between patients there was a decline in the range of BED values as the overall treatment time increased. This increased treatment time was partly a function of the slow decline in the activity of the sources with time but predominantly due to changes in the number of iso-centers used. Thus, variations in BED values did not correlate with prescription dose but was modified by the overall treatment time.
We propose a computational scheme for uncalibrated reconstruction of scene structure up to a relief transformation from binocular disparities. This scheme, which we call regional disparity correction (RDC), is motivated both by computational considerations and by psychophysical observations regarding human stereoscopic depth perception. We describe an implementation of RDC, and demonstrate its performance experimentally. As an example of applications of RDC, we show how it can be used to align a three-dimensional object model with an uncalibrated disparity field.
This paper addresses the problem of computing cues to the three-dimensional structure of surfaces in the world directly from the local structure of the brightness pattern of a binocular image pair. The geometric information content of the gradient of binocular disparity is analyzed for the general case of a fixating system with symmetric or asymmetric vergence, and with either known or unknown viewing geometry. A computationally inexpensive technique which exploits this analysis is proposed. This technique allows a local estimate of surface orientation to be computed directly from the local statistics of the left and right image brightness gradients, without iterations or search. The viability of the approach is demonstrated with experimental results for both synthetic and natural gray-level images.
Projective distortion of surface texture observed in a perspective image can provide direct information about the shape of the underlying surface. Previous theories have generally concerned planar surfaces; in this paper we present a systematic analysis of first- and second-order texture distortion cues for the case of a smooth curved surface. In particular, we analyze several kinds of texture gradients and relate them to surface orientation and surface curvature. The local estimates obtained from these cues can be integrated to obtain a global surface shape, and we show that the two surfaces resulting from the well-known tilt ambiguity in the local foreshortening cue typically have qualitatively different shapes. As an example of a practical application of the analysis, a shape from texture algorithm based on local orientation-selective filtering is described, and some experimental results are shown.
Rotationally symmetric operations in the image domain may give rise to shape distortions. This article describes a way of reducing this effect for a general class of methods for deriving 3-D shape cues from 2-D image data, which are based on the estimation of locally linearized distortion of brightness patterns. By extending the linear scale-space concept into an affine scale-space representation and performing affine shape adaption of the smoothing kernels, the accuracy of surface orientation estimates derived from texture and disparity cues can be improved by typically one order of magnitude. The reason for this is that the image descriptors, on which the methods are based, will be relative invariant under affine transformations, and the error will thus be confined to the higher-order terms in the locally linearized perspective mapping.
Gårding et al. (Vis Res 1995;35:703–722) proposed a two-stage theory of stereopsis. The first uses horizontal disparities for relief computations after they have been subjected to a process called disparity correction that utilises vertical disparities. The second stage, termed disparity normalisation, is concerned with computing metric representations from the output of stage one. It uses vertical disparities to a much lesser extent, if at all, for small field stimuli. We report two psychophysical experiments that tested whether human vision implements this two-stage theory. They tested the prediction that scaling vertical disparities to simulate different viewing distances to the fixation point should affect the perceived amplitudes of vertically but not horizontally oriented ridges. The first used elliptical half-cylinders and the ‘apparently circular cylinder' judgement task of Johnston (Vis Res 1991;31:1351–1360). The second experiment used parabolic ridges and the amplitude judgement task of Buckley and Frisby (Vis Res 1993;33:919–934). Both studies broadly confirmed the anisotropy prediction by finding that large scalings of vertical disparities simulating near distances had a strong effect on the perceived amplitudes of the vertically oriented stimuli but little effect on the horizontal ones. When distances >25 cm were simulated there were no significant differential effects and various methodological reasons are offered for this departure from expectations.
Gårding et al. (Vis Res 1995;35:703–722) proposed a two-stage theory of stereopsis. The first uses horizontal disparities for relief computations after they have been subjected to a process called disparity correction that utilises vertical disparities. The second stage, termed disparity normalisation, is concerned with computing metric representations from the output of stage one. It uses vertical disparities to a much lesser extent, if at all, for small field stimuli. We report two psychophysical experiments that tested whether human vision implements this two-stage theory. They tested the prediction that scaling vertical disparities to simulate different viewing distances to the fixation point should affect the perceived amplitudes of vertically but not horizontally oriented ridges. The first used elliptical half-cylinders and the ‘apparently circular cylinder’ judgement task of Johnston (Vis Res 1991;31:1351–1360). The second experiment used parabolic ridges and the amplitude judgement task of Buckley and Frisby (Vis Res 1993;33:919–934). Both studies broadly confirmed the anisotropy prediction by finding that large scalings of vertical disparities simulating near distances had a strong effect on the perceived amplitudes of the vertically oriented stimuli but little effect on the horizontal ones. When distances \25 cm were simulated there were no significant differential effects and various methodological reasons are offered for this departure from expectations. © 1998 Elsevier Science Ltd. All rights reserved.
This article describes a method for reducing the shape distortions due to scale-space smoothing that arise in the computation of 3-D shape cues using operators (derivatives) defined from scale-space representation. More precisely, we are concerned with a general class of methods for deriving 3-D shape cues from a 2-D image data based on the estimation of locally linearized deformations of brightness patterns. This class constitutes a common framework for describing several problems in computer vision (such as shape-from-texture, shape-from disparity-gradients, and motion estimation) and for expressing different algorithms in terms of similar types of visual front-end-operations. It is explained how surface orientation estimates will be biased due to the use of rotationally symmetric smoothing in the image domain. These effects can be reduced by extending the linear scale-space concept into an affine Gaussian scalespace representation and by performing affine shape adaptation of the smoothing kernels. This improves the accuracy of the surface orientation estimates, since the image descriptors, on which the methods are based, will be relative invariant under affine transformations, and the error thus confined to the higher-order terms in the locally linearized perspective transformation. A straightforward algorithm is presented for performing shape adaptation in practice. Experiments on real and synthetic images with known orientation demonstrate that in the presence of moderately high noise levels the accuracy is improved by typically one order of magnitude.
This paper addresses the problem of computing cues to the three-dimensional structure of surfaces in the world directly from the local structure of the brightness pattern of either a single monocular image or a binocular image pair.It is shown that starting from Gaussian derivatives of order up to two at a range of scales in scale-space, local estimates of (i) surface orientation from monocular texture foreshortening, (ii) surface orientation from monocular texture gradients, and (iii) surface orientation from the binocular disparity gradient can be computed without iteration or search, and by using essentially the same basic mechanism.The methodology is based on a multi-scale descriptor of image structure called the windowed second moment matrix, which is computed with adaptive selection of both scale levels and spatial positions. Notably, this descriptor comprises two scale parameters; a local scale parameter describing the amount of smoothing used in derivative computations, and an integration scale parameter determining over how large a region in space the statistics of regional descriptors is accumulated.Experimental results for both synthetic and natural images are presented, and the relation with models of biological vision is briefly discussed.
We propose an approach to determine the occurrence of low-parametric qualitative models from images by a hypothesis-and-test approach based on the coincidence of multiple cues, thereby avoiding complete reconstruction of the scene. A system is presented which applies the approach to finding instances of planar surfaces, as it is important in many tasks for mobile or manipulating robots. The system uses monocularly determined L-junctions and binocular disparities. A notable feature of the approach is that it finds the most conspicuous exemplar of the model first. This property seems quite relevant for an agent using vision to guide its behaviors, since the simplest solution becomes available early on.
A unified differential geometric framework for estimation of local surface shape and orientation from projective texture distortion is proposed, based on a differential version of the texture stationarity assumption introduced by Malik and Rosenholtz. This framework allows the information content of the gradient of any texture descriptor defined in a local coordinate frame to be characterized in a very compact form. The analysis encompasses both full affine texture descriptors and the classical "texture gradients". For estimation of local surface orientation and curvature from uncertain observations of affine texture distortion, the proposed framework allows the dimensionality of the search space to be reduced from five to one.<>
The pattern of retinal binocular disparities acquired by a fixating visual system depends on both the depth structure of the scene and the viewing geometry. This paper treats the problem of interpreting the disparity pattern in terms of scene structure without relying on estimates of fixation position from eye movement control and proprioception mechanisms. We propose a sequential decomposition of this interpretation process into disparity correction, which is used to compute three-dimensional structure up to a relief transformation, and disparity normalization, which is used to resolve the relief ambiguity to obtain metric structure. We point out that the disparity normalization stage can often be omitted, since relief transformations preserve important properties such as depth ordering and coplanarity. Based on this framework we analyse three previously proposed computational models of disparity processing; the Mayhew and Longuet-Higgins model, the deformation model and the polar angle disparity model. We show how these models are related, and argue that none of them can account satisfactorily for available psychophysical data. We therefore propose an alternative model, regional disparity correction. Using this model we derive predictions for a number of experiments based on vertical disparity manipulations, and compare them to available experimental data. The paper is concluded with a summary and a discussion of the possible architectures and mechanisms underling stereopsis in the human visual system.
This paper addresses the problem of computing cues to the three-dimensional structure of surfaces in the world directly from the local structure of the brightness pattern of either a single monocular image or a binocular image pair. It is shown that starting from Gaussian derivatives of order up to two at a range of scales in scale-space, local estimates of (i) surface orientation from monocular perspective texture foreshortening, (ii) surface orientation or curvature from monocular texture gradients, and (iii) surface orientation from the binocular disparity gradient can be computed without iteration or search, and by using essentially the same basic mechanism. The methodology is based on a multi-scale descriptor of image structure called the windowed second moment matrix, which is computed with adaptive selection of both scale levels and spatial positions. Notably, this descriptor comprises two scale parameters, a local scale describing the amount of smoothing used in derivative computations, and an integration scale determining over how large a region in space the statistics of local descriptors is accumulated. Experimental results for both synthetic and natural images are presented, and the relation with models of biological vision is brie y discussed. We would like to thank Jan-Olof Eklundh for continuous support and encouragement, as well as Narendra Ahuja at University of Illinois, and John P. Frisby at University of She eld for kindly providing several of the images used in the paper. This work was partially performed under the ESPRIT-BRA project INSIGHT. The support from the Swedish National Board for Industrial and Technical Development, NUTEK, is gratefully acknowledged. The rst author has carried out part of this work while visiting the AIVRU group at University of She eld, and he is grateful for their hospitality as well as for the nancial support of the Foundation Blance or Boncompagni-Ludovisi, n ee Bildt, and the Swedish Institute.
Witkin (1981) has proposed a maximum likelihood (ML) estimator of surface orientation based on the observed directional bias of projected texture elements. However, a drawback of this procedure is that the estimate is only defined indirectly in terms of a set of nonlinear equations. An alternative method is proposed, which allows an estimate of the surface orientation to be computed directly in a single step from certain simple statistics of the image data. We also show that this direct estimate allows Witkin's ML estimate to be computed to within 0.05 degrees in only two or three iterative steps. The performance of the new estimator is demonstrated experimentally and compared to that of the ML estimator, using both synthetic data and real gray-level images. >
A unified framework for shape from texture and contour is proposed. It is based on the assumption that the surface markings are not systematically compressed, or formally, that they are weakly isotropic.