Hazard models are the most commonly used tool to analyse time-to-event data. If more than one time scale is relevant for the event under study, models are required that can incorporate the dependence of a hazard along two (or more) time scales. Such models should be flexible to capture the joint influence of several times scales and nonparametric smoothing techniques are obvious candidates. P-splines offer a flexible way to specify such hazard surfaces, and estimation is achieved by maximizing a penalized Poisson likelihood. Standard observations schemes, such as right-censoring and left-truncation, can be accommodated in a straightforward manner. The model can be extended to proportional hazards regression with a baseline hazard varying over two scales. Generalized linear array model (GLAM) algorithms allow efficient computations, which are implemented in a companion R-package.
P-splines provide a flexible setting for modeling nonlinear model components based on a discretized penalty structure with a relatively simple computational backbone. Under a Bayesian inferential framework based on Markov chain Monte Carlo, estimates of model coefficients in P-splines models are typically obtained by means of Metropolis-type algorithms. These algorithms rely on a proposal distribution that has to be carefully chosen to generate Markov chains that efficiently explore the parameter space. To avoid such a sensitive tuning choice, we extend the Gibbs sampler to Bayesian P-splines models. In this model class, conditional posterior distributions of model coefficients are shown to have attractive mathematical properties. Taking advantage of these properties, we propose to sample the conditional posteriors by alternating between the adaptive rejection sampler when targets are log-concave and the Griddy-Gibbs sampler when targets are characterized by more complex shapes. The proposed Gibbs sampler for Bayesian P-splines (GSBPS) algorithm is shown to be an interesting tuning-free tool for inference in Bayesian P-splines models. Moreover, the GSBPS algorithm can be translated in a compact and user-friendly routine. After describing theoretical results, we illustrate the potential of our methodology in density estimation, Binomial regression, and smoothing of epidemic curves.
BACKGROUND:Low resolution nuclear magnetic resonance (LR-NMR) is a common technique to identify the constituents of complex materials (such as food and biological samples). The output of LR-NMR experiments is a relaxation signal which can be modelled as a type of convolution of an unknown density of relaxation times with decaying exponential functions, plus random Gaussian noise. The challenge is to estimate that density, a severely ill-posed problem. A complication is that non-negativity constraints need to be imposed in order to obtain valid results.SIGNIFICANCE AND NOVELTY:We present a smooth deconvolution model for solution of the inverse estimation problem in LR-NMR relaxometry experiments. We model the logarithm of the relaxation time density as a smooth function using (adaptive) P-splines while matching the expected residual magnetisations with the observed ones. The roughness penalty removes the singularity of the deconvolution problem, and the estimated density is positive by design (since we model its logarithm). The model is non-linear, but it can be linearized easily. The penalty has to be tuned for each given sample. We describe an efficient EM-type algorithm to optimize the smoothing parameter(s).RESULTS:We analyze a set of food samples (potato tubers). The relaxation spectra extracted using our method are similar to the ones described in the previous experiments but present sharper peaks. Using penalized signal regression we are able to accurately predict dry matter content of the samples using the estimated spectra as covariates.
Spatial transcriptomics is a novel technique that provides RNA-expression data with tissue-contextual annotations. Quality assessments of such techniques using end-user generated data are often lacking. Here, we evaluated data from the NanoString GeoMx Digital Spatial Profiling (DSP) platform and standard processing pipelines. We queried 72 ROIs from 12 glioma samples, performed replicate experiments of eight samples for validation, and evaluated five external datasets. The data consistently showed vastly different signal intensities between samples and experimental conditions that resulted in biased analysis. We evaluated the performance of alternative normalization strategies and show that quantile normalization can adequately address the technical issues related to the differences in data distributions. Compared to bulk RNA sequencing, NanoString DSP data show a limited dynamic range which underestimates differences between conditions. Weighted gene co-expression network analysis allowed extraction of gene signatures associated with tissue phenotypes from ROI annotations. Nanostring GeoMx DSP data therefore require alternative normalization methods and analysis pipelines.
Supplementary Table from Progression and Tumor Heterogeneity Analysis in Early Rectal Cancer
Supplementary Table 4 - PDF file 69K, CpGs associated with benefit from adjuvant PCV chemotherapy
We present a model for direct semi-parametric estimation of the State Price Density (SPD) implied in quoted option prices. We treat the observed prices as expected values of possible pay-offs at maturity, weighted by the unknown probability density function. We model the logarithm of the latter as a smooth function while matching the expected values of the potential pay-offs with the observed prices. This leads to a special case of the penalized composite link model. Our estimates do not rely on any parametric assumption on the underlying asset price dynamics and are consistent with no-arbitrage conditions. The model shows excellent performance in simulations and in application to real data.
Background: Spatially resolved transcriptomics is a novel and already highly recognized method that allows RNA sequencing results to be annotated with local tissue phenotypes. The NanoString GeoMx Digital Spatial Profiling (DSP) Platform allows users to collect RNA expression data from manually selected Regions of Interest (ROIs) on FFPE tissue sections. Here, we extensively evaluated data from the DSP platform with its associated pipeline and identify significant background noise interference issues which compromise data interpretation. Alternative and more suitable workflows are presented for correct data analysis. Methods: In this study, 12 paired tumor samples were collected from six glioma patients who underwent two separate resections. For all patients, the first resection was a low grade astrocytoma (WHO grade II or III) and the second resection was a high grade astrocytoma (WHO grade IV). The DSP platform was used to collect expression data of 1,800 genes from 72 ROIs (i.e. 6 per sample). Biological replicates were made of eight tumors from four patients. Gene expression data was normalized with both standard NanoString methods and several alternative methods (e.g. DeSeq2, gamma fit correction and quantile normalization). Weighted Gene Co-expression Network analysis (WGCNA) was used for biological validation. In addition to our own study, six publicly available NanoString DSP datasets were evaluated. Results: Data distributions of all glioma samples, when exposed to standard data processing, were burdened with significant background noise interference. Notably, differences in noise interference were largest between biologically distinct tumor subgroups (i.e. between first and second glioma resections), which was confirmed in replicate experiments. The noise interference patterns were also present in all six publicly available NanoString DSP datasets which will invariably lead to incorrect interpretation of the underlying biology. To correct for noise interference, we tested several normalization methods. The relatively crude quantile normalization method provided the least biased result and showed the highest concordance with bulk RNA sequencing data. To evaluate the biological validity of our alternative approach, we used T cell counts from our tissue regions as an independent parameter, that were quantified using immune fluorescence. Unsupervised WGCNA identified gene clusters enriched for lymphocyte genes that highly correlated with T cell quantities in ROIs, confirming that alternative normalization can extract a biological signal from the DSP platform. Conclusion: The DSP Platform platform suffers from significant noise interference when using standard analysis tools that obscure its results. Here, we revised the workflow and provide an alternative normalization that adequately addresses noise interference and enables correct interpretation of gene expression data. Citation Format: Levi van Hijfte, Marjolein Geurts, Wies R. Vallentgoed, Paul H. Eilers, Peter A. Sillevis Smitt, Reno Debets, Pim J. French. Spatial transcriptomics: Data processing revisited to address noise interference [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 1228.
We present a fast and simple algorithm for super-resolution with single images. It is based on penalized least squares regression and exploits the tensor structure of two-dimensional convolution. A ridge penalty and a difference penalty are combined; the former removes singularities, while the latter eliminates ringing. We exploit the conjugate gradient algorithm to avoid explicit matrix inversion. Large images are handled with ease: zooming a 100 by 100 pixel image to 800 by 800 pixels takes less than a second on an average PC. Several examples, from applications in wide-field fluorescence microscopy, illustrate performance.
On May 25, 2022, Iain David Currie died from cancer. Iain is survived by his children Janet and Colin and two grandchildren. He had lost his beloved wife Elizabeth in 2020. Iain was one of the gentlest people in Statistics. He spoke softly, but always had a clear opinion. His thinking was very sharp and precise. We both have enjoyed working with him for decades. It was always stimulating and rewarding, as it led to a series of well-received papers. Iain's early work was mainly on details of classical statistics, studying distributions and testing problems, with an occasional paper on agricultural trials. This changed after Maria became his PhD student in 1995. They developed additive models, using local linear regression for smoothing. As a supervisor, he was thorough and precise in his work, curious and down to earth. His comments were always insightful and he enjoyed challenging everyone else's ideas, but he did all this with a great smile and a very Scottish sense of humour. He was a great teacher, supervisor and mentor: patient, always keen to help and give advice, extremely generous with his knowledge, and always positive and encouraging. Since 1999 Iain was a regular participant in the International Workshops on Statistical Modelling (IWSM). The previous year Maria had attended the Workshop in New Orleans, and her feedback convinced Iain that he should join in too. This almost certainly changed his life. At the IWSM in 2000, held in Bilbao, Iain contacted Paul Eilers about P-splines, a combination of B-splines and discrete penalties. Iain was particularly interested in two-dimensional applications, for agricultural trials and mortality tables. This was the start of a very fruitful cooperation between the three of us, resulting in several papers that found a wide audience. Of special interest is the generalised linear array model, a very effective algorithm for smoothing of multidimensional grids. Iain's mastery of algebra was of crucial importance for its development. Iain was a central figure in a series of ‘P-splines fiestas’, that mainly took place in Madrid and Bilbao. A group of P-splines aficionados had grown over the years, with a majority of statisticians that had been supervised as PhD students by Maria. This informal group developed many new ideas and put them into papers and software. The discussions could be fierce at times, but afterwards we enjoyed the tapas or the pintxos with a good glass of beer or wine in a very good mood. Iain had a special talent for explaining complicated subjects in a very precise and clear way. Of special interest are his blogs on smoothing of mortality tables, on the Website of Longevitas, a company for which he was a consultant. Many of his ideas can be found in the book Modelling Mortality with Actuarial Applications (Cambridge 2018), that he wrote with Angus MacDonald and Stephen Richards. Iain's disease had robbed him of his old energy, but his mind was sharp as ever. When Paul visited him in February 2022, Iain showed him some of his latest ideas on constrained estimation and identifiability. Initially he was reluctant, but finally he got convinced that it was a good idea to make a complete paper out of it. That was planned for a special issue of Statistical Modelling, dedicated to Brian Marx, who died in November 2021. Iain finished the paper quickly and submitted an excellent manuscript near the end of March. To cite his own words: ‘Finishing the paper was a triumph whatever its quality may be. I can still scarcely believe it. One last hurrah!’ Exactly 2 months later we lost him forever. Iain was a wonderful friend. He has contributed a lot to the field of statistics. It is visible in his papers, but there is a larger invisible part: the many things we learned from Iain. We owe him very much and we will never forget him.
The receptive field (RF) of a visual neuron is the region of the space that elicits neuronal responses. It can be mapped using different techniques that allow inferring its spatial and temporal properties. Raw RF maps (RFmaps) are usually noisy, making it difficult to obtain and study important features of the RF. A possible solution is to smooth them using P-splines. Yet, raw RFmaps are characterized by sharp transitions in both space and time. Their analysis thus asks for spatiotemporal adaptive P-spline models, where smoothness can be locally adapted to the data. However, the literature lacks proposals for adaptive P-splines in more than two dimensions. Furthermore, the extra flexibility afforded by adaptive P-spline models is obtained at the cost of a high computational burden, especially in a multidimensional setting. To fill these gaps, this work presents a novel anisotropic locally adaptive P-spline model in two (e.g., space) and three (space and time) dimensions. Estimation is based on the recently proposed SOP (Separation of Overlapping Precision matrices) method, which provides the speed we look for. Besides the spatiotemporal analysis of the neuronal activity data that motivated this work, the practical performance of the proposal is evaluated through simulations, and comparisons with alternative methods are reported.
Array Algorithms are defined as functional algorithms where each step of the algorithm results in a function being applied to an array, producing an array result. Array algorithms are compared with non-array algorithms. A brief rationale for teaching array algorithms is given together with an example which shows that array algorithms sometimes lead to surprising results.
Sub-diffraction or super-resolution fluorescence imaging allows the visualization of the cellular morphology and interactions at the nanoscale. Statistical analysis methods such as super-resolution optical fluctuation imaging (SOFI) obtain an improved spatial resolution by analyzing fluorophore blinking but can be perturbed by the presence of non-stationary processes such as photodestruction or fluctuations in the illumination. In this work, we propose to use Whittaker smoothing to remove these smooth signal trends and retain only the information associated to independent blinking of the emitters, thus enhancing the SOFI signals. We find that our method works well to correct photodestruction, especially when it occurs quickly. The resulting images show a much higher contrast, strongly suppressed background and a more detailed visualization of cellular structures. Our method is parameter-free and computationally efficient, and can be readily applied on both two-dimensional and three-dimensional data.