Background: Cervical cancer develops over several years; screening and early diagnosis have decreased the incidence and mortality threefold over the last fifty years. Opportunities for the application of imaging and automation in the screening process exist in settings where resources are limited. Methods: Patients with high-grade squamous intraepithelial lesions (SIL) underwent imaging with a Multispectral Digital Colposcopy (MDC) prior to have a loop excision of the cervix. The image taken with white light was annotated by a clinician. The excised specimen was mapped by the study histopathologist blinded to the MDC data. This map was used to define areas of high grade in the excised tissue. Eleven reviewers mapped the histopathologic data into the MDC images. The reviewers' maps were analyzed and areas of agreement were calculated. We compared the result of a boosted tree classifier with a previously developed ensemble classifier. Results: Using a boosted tree classifier we obtained a sensitivity of 95%, a specificity of 96%, and an accuracy of 96% on the training sets. When we applied the classifier to a test set, we obtained a sensitivity of 82%, a specificity of 81%, and an accuracy of 81%. The boosted tree classifier performed better than the previously developed ensemble classifier. Conclusion: Here we presented promising results which show that a boosted tree analysis on MDC images is a method that could be used as an adjunct to colposcopy and would result in greater diagnostic accuracy compared to existing methods.
Functional data are defined as realizations of random functions (mostly smooth functions) varying over a continuum, which are usually collected on discretized grids with measurement errors. In order to accurately smooth noisy functional observations and deal with the issue of high-dimensional observation grids, we propose a novel Bayesian method based on the Bayesian hierarchical model with a Gaussian-Wishart process prior and basis function representations. We first derive an induced model for the basis-function coefficients of the functional data, and then use this model to conduct posterior inference through Markov chain Monte Carlo methods. Compared to the standard Bayesian inference that suffers serious computational burden and instability in analyzing high-dimensional functional data, our method greatly improves the computational scalability and stability, while inheriting the advantage of simultaneously smoothing raw observations and estimating the mean-covariance functions in a nonparametric way. In addition, our method can naturally handle functional data observed on random or uncommon grids. Simulation and real studies demonstrate that our method produces similar results to those obtainable by the standard Bayesian inference with low-dimensional common grids, while efficiently smoothing and estimating functional data with random and high-dimensional observation grids when the standard Bayesian inference fails. In conclusion, our method can efficiently smooth and estimate high-dimensional functional data, providing one way to resolve the curse of dimensionality for Bayesian functional data analysis with Gaussian-Wishart processes.
Many scientific studies measure different types of high-dimensional signals or images from the same subject, producing multivariate functional data. These functional measurements carry different types of information about the scientific process, and a joint analysis that integrates information across them may provide new insights into the underlying mechanism for the phenomenon under study. Motivated by fluorescence spectroscopy data in a cervical pre-cancer study, a multivariate functional response regression model is proposed, which treats multivariate functional observations as responses and a common set of covariates as predictors. This novel modeling framework simultaneously accounts for correlations between functional variables and potential multi-level structures in data that are induced by experimental design. The model is fitted by performing a two-stage linear transformation—a basis expansion to each functional variable followed by principal component analysis for the concatenated basis coefficients. This transformation effectively reduces the intra- and inter-function correlations and facilitates fast and convenient calculation. A fully Bayesian approach is adopted to sample the model parameters in the transformed space, and posterior inference is performed after inverse-transforming the regression coefficients back to the original data domain. The proposed approach produces functional tests that flag local regions on the functional effects, while controlling the overall experiment-wise error rate or false discovery rate. It also enables functional discriminant analysis through posterior predictive calculation. Analysis of the fluorescence spectroscopy data reveals local regions with differential expressions across the pre-cancer and normal samples. These regions may serve as biomarkers for prognosis and disease assessment.
In the cervix in vivo multi-excitation multi-emission (AFI) data can have high sensitivity (~90% ) but specificity is affected by confounding tissue structures. OCT-AFI combined imaging has potential to identify these confounding structures.
Functional data, with basic observational units being functions (e.g., curves, surfaces) varying over a continuum, are frequently encountered in various applications. While many statistical tools have been developed for functional data analysis, the issue of smoothing all functional observations simultaneously is less studied. Existing methods often focus on smoothing each individual function separately, at the risk of removing important systematic patterns common across functions. We propose a nonparametric Bayesian approach to smooth all functional observations simultaneously and nonparametrically. In the proposed approach, we assume that the functional observations are independent Gaussian processes subject to a common level of measurement errors, enabling the borrowing of strength across all observations. Unlike most Gaussian process regression models that rely on pre-specified structures for the covariance kernel, we adopt a hierarchical framework by assuming a Gaussian process prior for the mean function and an Inverse-Wishart process prior for the covariance function. These prior assumptions induce an automatic mean-covariance estimation in the posterior inference in addition to the simultaneous smoothing of all observations. Such a hierarchical framework is flexible enough to incorporate functional data with different characteristics, including data measured on either common or uncommon grids, and data with either stationary or nonstationary covariance structures. Simulations and real data analysis demonstrate that, in comparison with alternative methods, the proposed Bayesian approach achieves better smoothing accuracy and comparable mean-covariance estimation results. Furthermore, it can successfully retain the systematic patterns in the functional observations that are usually neglected by the existing functional data analyses based on individual-curve smoothing.
Introduction Since colposcopy helps to detect cervical cancer in its precancerous stages, as new strategies and technologies are developed for the clinical management of cervical neoplasia, precisely determining the accuracy of colposcopy is important for characterizing its continued role. Our objective was to employ a more precise methodology to estimate of the accuracy of colposcopy to better reflect clinical practice. Study design For each patient, we compared the worst histology result among colposcopically positive sites to the worst histology result among all sites biopsied, thereby more accurately determining the number of patients that would have been underdiagnosed by colposcopy than previously estimated. Materials and Methods We utilized data from a clinical trial in which 850 diagnostic patients had been enrolled. Seven hundred and ninety-eight of the 850 patients had been examined by colposcopy, and biopsy samples were taken at colposcopically normal and abnormal sites. Our endpoints of interest were the percentages of patients underdiagnosed, and sensitivity and specificity of colposcopy. Results With the threshold of low-grade squamous intraepithelial lesions for positive colposcopy and histology diagnoses, the sensitivity of colposcopy decreased from our previous assessment of 87.0% to 74.0%, while specificity remained the same. The drop in sensitivity was the result of histologically positive sites that were diagnosed as negative by colposcopy. Thus, 28.4% of the 798 patients in this diagnostic group would have had their condition underdiagnosed by colposcopy in the clinic. Conclusions In utilizing biopsies at multiple sites of the cervix, we present a more precise methodology for determining the accuracy of colposcopy. The true accuracy of colposcopy is lower than previously estimated. Nevertheless, our results reinforce previous conclusions that colposcopy has an important role in the diagnosis of cervical precancer.
We consider testing equality of mean functions from two samples of functional data. A novel test based on the adaptive Neyman methodology applied to the Hotelling’s T-squared statistic is proposed. Under the enlarged null hypothesis that the distributions of the two populations are the same, randomization methods are proposed to find a null distribution which gives accurate significance levels. An extensive simulation study is presented which shows that the proposed test works very well in comparison with several other methods under a variety of alternatives and is one of the best methods for all alternatives, whereas the other methods all show weak power at some alternatives. An application to a real-world data set demonstrates the applicability of the method.
Although the Papanicolaou smear has been successful in decreasing cervical cancer incidence in the developed world, there exist many challenges for implementation in the developing world. Quantitative cytology, a semi-automated method that quantifies cellular image features, is a promising screening test candidate. The nested structure of its data (measurements of multiple cells within a patient) provides challenges to the usual classification problem. Here we perform a comparative study of three main approaches for problems with this general data structure: a) extract patient-level features from the cell-level data; b) use a statistical model that accounts for the hierarchical data structure; and c) classify at the cellular level and use an ad hoc approach to classify at the patient level. We apply these methods to a dataset of 1,728 patients, with an average of 2,600 cells collected per patient and 133 features measured per cell, predicting whether a patient had a positive biopsy result. The best approach we found was to classify at the cellular level and count the number of cells that had a posterior probability greater than a threshold value, with estimated 61% sensitivity and 89% specificity on independent data. Recent statistical learning developments allowed us to achieve high accuracy.
We are investigating spectroscopic devices designed to make in vivo cervical tissue measurements to detect pre-cancerous and cancerous lesions.All devices have the same design and ideally should record identical measurements.However, we observed consistent differences among them.An experiment was designed to study the sources of variation in the measurements recorded.Here we present a log additive statistical model that incorporates the sources of variability we identified.Based on this model, we estimated correction factors from the experimental data needed to eliminate the inter-device variability and other sources of variation.These correction factors are intended to improve the accuracy and repeatability of such devices when making future measurements on patient tissue.
We report clinical imaging results for an MDC that has been applied to more than 50 patients with suspected cervical precancer and analyze against matched histopathology results. For high grade dysplasia: sensitivity ~90%, specificity ~55%.
This article presents an empirical Bayesian code tuning method based on a Gaussian process model for estimating adjustable theory parameters in a complex computer simulation code by using both computer simulation data and real experimental data. Some parameters of the metamodel are estimated from the data by the maximum likelihood method, and those estimates are then used to obtain the maximum a posterior estimate of theory parameters. Four transport parameters of the theoretical nuclear fusion model are estimated by applying this method to computational nuclear fusion devices (tokamak). The approximated standard errors of estimates are obtained by using the Fisher information matrix. The posterior probability of a parameter is computed to test a hypothesis about the parameter.
Optical spectroscopy has been proposed as an accurate and low-cost alternative for detection of cervical intraepithelial neoplasia. We previously published an algorithm using optical spectroscopy as an adjunct to colposcopy and found good accuracy (sensitivity=1.00 [95% confidence interval (CI)=0.92 to 1.00], specificity=0.71 [95% CI=0.62 to 0.79]). Those results used measurements taken by expert colposcopists as well as the colposcopy diagnosis. In this study, we trained and tested an algorithm for the detection of cervical intraepithelial neoplasia (i.e., identifying those patients who had histology reading CIN 2 or worse) that did not include the colposcopic diagnosis. Furthermore, we explored the interaction between spectroscopy and colposcopy, examining the importance of probe placement expertise. The colposcopic diagnosis-independent spectroscopy algorithm had a sensitivity of 0.98 (95% CI=0.89 to 1.00) and a specificity of 0.62 (95% CI=0.52 to 0.71). The difference in the partial area under the ROC curves between spectroscopy with and without the colposcopic diagnosis was statistically significant at the patient level (p=0.05) but not the site level (p=0.13). The results suggest that the device has high accuracy over a wide range of provider accuracy and hence could plausibly be implemented by providers with limited training.
A new hybrid simulation method is introduced that adaptively combines the four most efficient simulation methods at different scales by dynamically partitioning the reactions according to both the reaction propensity and the molecular abundance. Simulation results prove it to be accurate but much more efficient than SSA. This could provide an accurate and efficient multiscale solution for biochemical reaction simulations in systems biology, since it covers more scales efficiently and can be easily updated by “plugging-in” the latest algorithm in a specified level.
There is an urgent global need for effective and affordable approaches to cervical cancer screening and diagnosis. In developing nations, cervical malignancies remain the leading cause of cancer-related deaths in women. This reality may be difficult to accept given that these deaths are largely preventable; where cervical screening programs have been implemented, cervical cancer–related deaths have decreased dramatically. In developed countries, the challenges of cervical disease stem from high costs and overtreatment. The National Cancer Institute–funded Program Project is evaluating the applicability of optical technologies in cervical cancer. The mandate of the project is to create tools for disease detection and diagnosis that are inexpensive, require minimal expertise, are more accurate than existing modalities, and can be feasibly implemented in a variety of clinical settings. This article presents the status and long-term goals of the project.