We make an observation that facilitates exact likelihood-based inference for the parameters of the popular ARFIMA model without requiring stationarity by allowing the upper bound d for the memory parameter d to exceed 0.5: estimating the parameters of a single nonstationary ARFIMA model is equivalent to estimating the parameters of a sequence of stationary ARFIMA models. This allows for the use of existing methods for evaluating the likelihood for an invertible and stationary ARFIMA model. This enables improved inference because many standard methods perform poorly when estimates are close to the boundary of the parameter space. It also allows us to leverage the wealth of likelihood approximations that have been introduced for estimating the parameters of a stationary process. We explore how estimation of the memory parameter d depends on the upper bound d and introduce adaptive procedures for choosing d. We show via simulation how our adaptive procedures estimate the memory parameter well, relative to existing alternatives, when the true value is as large as 2.5.
The ability to interpret machine learning models has become increasingly important as their usage in data science continues to rise. Most current interpretability methods are optimized to work on either (i) a global scale, where the goal is to rank features based on their contributions to overall variation in an observed population, or (ii) the local level, which aims to detail on how important a feature is to a particular individual in the data set. In this work, a new operator is proposed called the “GlObal And Local Score” (GOALS): a simple post hoc approach to simultaneously assess local and global feature variable importance in nonlinear models. Motivated by problems in biomedicine, the approach is demonstrated using Gaussian process regression where the task of understanding how genetic markers are associated with disease progression both within individuals and across populations is of high interest. Detailed simulations and real data analyses illustrate the flexible and efficient utility of GOALS over state-of-the-art variable importance strategies.
In many regression settings the unknown coefficients may have some known structure, for instance they may be ordered in space or correspond to a vectorized matrix or tensor. At the same time, the unknown coefficients may be sparse, with many nearly or exactly equal to zero. However, many commonly used priors and corresponding penalties for coefficients do not encourage simultaneously structured and sparse estimates. In this article we develop structured shrinkage priors that generalize multivariate normal, Laplace, exponential power and normal-gamma priors. These priors allow the regression coefficients to be correlated a priori without sacrificing elementwise sparsity or shrinkage. The primary challenges in working with these structured shrinkage priors are computational, as the corresponding penalties are intractable integrals and the full conditional distributions that are needed to approximate the posterior mode or simulate from the posterior distribution may be nonstandard. We overcome these issues using a flexible elliptical slice sampling procedure, and demonstrate that these priors can be used to introduce structure while preserving sparsity. Supplementary materials for this article are available online.
Pathwise coordinate descent algorithms have been used to compute entire solution paths for lasso and other penalized regression problems quickly with great success. They improve upon cold start algorithms by solving the problems that make up the solution path sequentially for an ordered set of tuning parameter values, instead of solving each problem separately. However, extending pathwise coordinate descent algorithms to more the general bridge or power family of lq penalties is challenging. Faster algorithms for computing solution paths for these penalties are needed because lq penalized regression problems can be nonconvex and especially burdensome to solve. In this article, we show that a reparameterization of lq penalized regression problems is more amenable to pathwise coordinate descent algorithms. This allows us to improve computation of the mode-thresholding function for lq penalized regression problems in practice and introduce two separate pathwise algorithms. We show that either pathwise algorithm is faster than the corresponding cold start alternative, and demonstrate that different pathwise algorithms may be more likely to reach better solutions. Supplemental materials for this article are available online.
Abstract Background: Non-small cell lung cancer (NSCLC) survival is closely linked to metastatic progression. 5-year survival rates drop from >65% when localized to ~9% when distant metastases occur. Despite this impact on mortality, drivers of overall metastasis and site-specific organotropism are not well understood in NSCLC. To investigate these questions, we focused on Lung adenocarcinoma (LUAD), the most common subtype of NSCLC. Although previous efforts have been made to examine LUAD organotropism, we sought to expand upon longitudinal modeling approaches and explore associations in an independent large clinical and genomic annotated cohort. Methods: Using the artificial intelligence model from (Kehl et al. 2021), we annotated a clinico-genomic cohort of 2777 patients with lung adenocarcinoma, resulting in 59,177 annotated imaging reports. This resulted in time to metastatic site annotations for each patient, with median of 2 years follow up. Metastatic sites included brain, bone, adrenal, liver, lymph node, and mesentery. All patients were evaluated at the Dana-Farber Cancer Institute, and each patient had at least one tumor biopsy sequenced with targeted panel sequencing of 400+ cancer related genes. To infer genomic and clinical correlates of site-specific metastasis, we performed cause-specific and Fine-Gray hazard modeling in this cohort. We used the recurrent event Andersen-Gill model to infer predictors of longitudinal metastasis rates. Results: Site-specific analyses yielded genes whose mutant status were significantly associated with a change in risk and/or rate of certain metastatic sites. We found that NOTCH1 is significantly associated with an increased rate and risk of brain metastases (HR=2.46, P<0.001). SETD2 is associated with decreased incidence of brain (HR=0.42, P<0.001). SMARCA4 is significantly associated with increased incidence and rate of adrenal metastasis (HR=1.64, P<0.01), while NF1 is significantly associated with decreased rate of adrenal metastasis (HR=0.72, P<0.01). In a multivariable model, Female patients and younger patients had a lower rate of new metastatic sites. As previously seen, TP53 and SMARCA4 mutations associated with higher rates of new metastatic sites (HR=1.17, P<0.001; HR=1.36, P<0.001 respectively), and SETD2 is associated with a lower rate of new metastatic sites (HR=0.71, P<0.001). Conclusions: By using time-to-event analysis with longitudinal metastatic annotations, we identify potential drivers of overall metastasis and site-specific organotropism in LUAD patients. Identification of new associations and concordance of our results with previous studies also supports the utility of artificial-intelligence-annotated datasets, allowing for larger cohorts without the need for manual annotation. Citation Format: Tyler Aprati, Michael Manos, Giuseppe Tarantino, Marc Glettig, Maryclare Griffin, Alexander Gusev, Kenneth Kehl, David Liu. Leveraging machine-learning approaches to dissect drivers of clinical metastatic dynamics in lung adenocarcinoma [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 2750.
Lightning is a destructive and highly visible product of severe storms, yet there is still much to be learned about the conditions under which lightning is most likely to occur. The GOES-16 and GOES-17 satellites, launched in 2016 and 2018 by NOAA and NASA, collect a wealth of data regarding individual lightning strike occurrence and potentially related atmospheric variables. The acute nature and inherent spatial correlation in lightning data renders standard regression analyses inappropriate. Further, computational considerations are foregrounded by the desire to analyze the immense and rapidly increasing volume of lightning data. We present a new computationally feasible method that combines spectral and Laplace approximations in an EM algorithm, denoted SLEM, to fit the widely popular log-Gaussian Cox process model to large spatial point pattern datasets. In simulations, we find SLEM is competitive with contemporary techniques in terms of speed and accuracy. When applied to two lightning datasets, SLEM provides better out-of-sample prediction scores and quicker runtimes, suggesting its particular usefulness for analyzing lightning data, which tend to have sparse signals.
Measurements of many biological processes are characterized by an initial trend period followed by an equilibrium period. Scientists may wish to quantify features of the two periods, as well as the timing of the change point. Specifically, we are motivated by problems in the study of electrical cell-substrate impedance sensing (ECIS) data. ECIS is a popular new technology which measures cell behavior non-invasively. Previous studies using ECIS data have found that different cell types can be classified by their equilibrium behavior. However, it can be challenging to identify when equilibrium has been reached, and to quantify the relevant features of cells' equilibrium behavior. In this paper, we assume that measurements during the trend period are independent deviations from a smooth nonlinear function of time, and that measurements during the equilibrium period are characterized by a simple long memory model. We propose a method to simultaneously estimate the parameters of the trend and equilibrium processes and locate the change point between the two. We find that this method performs well in simulations and in practice. When applied to ECIS data, it produces estimates of change points and measures of cell equilibrium behavior which offer improved classification of infected and uninfected cells.
Many penalized maximum likelihood estimators correspond to posterior mode estimators under specific prior distributions. Appropriateness of a particular class of penalty functions can therefore be interpreted as the appropriateness of a prior for the parameters. For example, the appropriateness of a lasso penalty for regression coefficients depends on the extent to which the empirical distribution of the regression coefficients resembles a Laplace distribution. We give a testing procedure of whether or not a Laplace prior is appropriate and accordingly, whether or not using a lasso penalized estimate is appropriate. This testing procedure is designed to have power against exponential power priors which correspond to ℓ_q penalties. Via simulations, we show that this testing procedure achieves the desired level and has enough power to detect violations of the Laplace assumption when the numbers of observations and unknown regression coefficients are large. We then introduce an adaptive procedure that chooses a more appropriate prior and corresponding penalty from the class of exponential power priors when the null hypothesis is rejected. We show that this can improve estimation of the regression coefficients both when they are drawn from an exponential power distribution and when they are drawn from a spike-and-slab distribution.
Consider the problem of estimating the entries of an unknown mean matrix or tensor given a single noisy realization. In the matrix case, this problem can be addressed by decomposing the mean matrix into a component that is additive in the rows and columns, i.e. the additive ANOVA decomposition of the mean matrix, plus a matrix of elementwise effects, and assuming that the elementwise effects may be sparse. Accordingly, the mean matrix can be estimated by solving a penalized regression problem, applying a lasso penalty to the elementwise effects. Although solving this penalized regression problem is straightforward, specifying appropriate values of the penalty parameters is not. Leveraging the posterior mode interpretation of the penalized regression problem, moment-based empirical Bayes estimators of the penalty parameters can be defined. Estimation of the mean matrix using these moment-based empirical Bayes estimators can be called LANOVA penalization, and the corresponding estimate of the mean matrix can be called the LANOVA estimate. The empirical Bayes estimators are shown to be consistent. Additionally, LANOVA penalization is extended to accommodate sparsity of row and column effects and to estimate an unknown mean tensor. The behavior of the LANOVA estimate is examined under misspecification of the distribution of the elementwise effects, and LANOVA penalization is applied to several datasets, including a matrix of microarray data, a three-way tensor of fMRI data and a three-way tensor of wheat infection data.
Respondent-driven sampling (RDS) is a method for sampling from a target population by leveraging social connections. RDS is invaluable to the study of hard-to-reach populations. However, RDS is costly and can be infeasible. RDS is infeasible when RDS point estimators have small effective sample sizes (large design effects) or when RDS interval estimators have large confidence intervals relative to estimates obtained in previous studies or poor coverage. As a result, researchers need tools to assess whether or not estimation of certain characteristics of interest for specific populations is feasible in advance. In this paper, we develop a simulation-based framework for using pilot data-in the form of a convenience sample of aggregated, egocentric data and estimates of subpopulation sizes within the target population-to assess whether or not RDS is feasible for estimating characteristics of a target population. in doing so, we assume that more is known about egos than alters in the pilot data, which is often the case with aggregated, egocentric data in practice. We build on existing methods for estimating the structure of social networks from aggregated, egocentric sample data and estimates of subpopulation sizes within the target population. We apply this framework to assess the feasibility of estimating the proportion male, proportion bisexual, proportion depressed and proportion infected with HIV/AIDS within three spatially distinct target populations of older lesbian, gay and bisexual adults using pilot data from the caring and Aging with Pride Study and the Gallup Daily Tracking Survey. We conclude that using an RDS sample of 300 subjects is infeasible for estimating the proportion male, but feasible for estimating the proportion bisexual, proportion depressed and proportion infected with HIV/AIDS in all three target populations.
The current bioassay development literature lacks the use of statistically robust methods for calculating the limit of detection of a given assay. Instead, researchers often employ simple methods that provide a rough estimate of the limit of detection, often without a measure of the confidence in the estimate. This scarcity of robust methods is likely due to a realistic preference for simple and accessible methods and to a lack of such methods that have reduced the concepts of limit of detection theory to practice for the specific application of bioassays. Here, we have developed a method for determining limits of detection for bioassays that is statistically robust and reduced to practice in a clear and accessible manner geared at researchers, not statisticians. This method utilizes a four-parameter logistic curve fit to translate signal intensity to analyte concentration, which is a curve that is commonly employed in quantitative bioassays. This method generates a 95% confidence interval of the limit of detection estimate to provide a measure of uncertainty and a means by which to compare the analytical sensitivities of different assays statistically. We have demonstrated this method using real data from the development of a paper-based influenza assay in our laboratory to illustrate the steps and features of the method. Using this method, assay developers can calculate statistically valid limits of detection and compare these values for different assays to determine when a change to the assay design results in a statistically significant improvement in analytical sensitivity.