Accurate diagnostic tests are essential for effective screening and treatment. However, individual biomarkers often fail to provide sufficient diagnostic accuracy, as they typically capture only one aspect of the complex disease process. Combining multiple biomarkers, each capturing a distinct mechanism, can help constructing more informative diagnostic tests. In practice, logistic regression is used as the default to combine biomarkers, but it can perform poorly when biomarker distributions exhibit skewness or differ across disease groups. Nonparametric methods provide more flexibility but generally require large sample sizes that are infrequently available in biomedical research. We propose a novel framework called transformation discriminant analysis which combines biomarkers through the likelihood ratio function to construct theoretically optimal diagnostic scores. Transformation discriminant analysis balances between flexibility and efficiency. It can accommodate a wide range of distributional shapes and disease-specific dependence structures while remaining fully parametric. This allows for likelihood inference and strong performance even in small-sample settings. We evaluate TDA through simulations and benchmark its performance against commonly used methods. Finally, we illustrate its utility in constructing an optimal diagnostic test for hepatocellular carcinoma, a disease with no single ideal biomarker. An open-source R implementation is provided for reproducibility and broader application.
This rejoinder addresses the three discussions of our manuscript on "Nonparanormal Adjusted Marginal Inference". We appreciate the thoughtful assessments and provide clarifications and refinements in response to the points raised.
Identifying reliable biomarkers for predicting clinical events in longitudinal studies is important for accurate disease prognosis and for guiding development of new treatments. However, prognostic studies are often observational, making it difficult to account for patient heterogeneity. In amyotrophic lateral sclerosis (ALS), factors such as age, site of onset and genetic status influence both survival and biomarker levels, yet their impact on the prognostic accuracy of biomarkers over time remains unclear. While time-dependent receiver operating characteristic methods have been developed to handle censored time-to-event outcomes, most do not adjust for covariates. To address this, we propose the nonparanormal prognostic biomarker framework, which models the joint distribution of the biomarker and event time while accounting for covariates. This allows estimation of covariate-specific time-dependent receiver operating characteristic curves and related summary measures. We apply the NPB framework to evaluate serum neurofilament light as a prognostic biomarker in ALS, showing that its accuracy varies over time and with patient characteristics. By capturing these covariate-specific effects, the NPB framework supports more targeted risk stratification and can potentially improve the design of clinical trials for new ALS treatments.
Ensembles improve prediction performance and allow uncertainty quantification by aggregating predictions from multiple models. In deep ensembling, the individual models are usually black box neural networks, or recently, partially interpretable semi-structured deep transformation models. However, interpretability of the ensemble members is generally lost upon aggregation. This is a crucial drawback of deep ensembles in high-stake decision fields, in which interpretable models are desired. We propose a novel transformation ensemble which aggregates probabilistic predictions with the guarantee to preserve interpretability and yield uniformly better predictions than the ensemble members on average. Transformation ensembles are tailored towards interpretable deep transformation models but are applicable to a wider range of probabilistic neural networks. In experiments on several publicly available data sets, we demonstrate that transformation ensembles perform on par with classical deep ensembles in terms of prediction performance, discrimination, and calibration. In addition, we demonstrate how transformation ensembles quantify both aleatoric and epistemic uncertainty, and produce minimax optimal predictions under certain conditions.
Although treatment effects can be estimated from observed outcome distributions obtained from proper randomization in clinical trials, covariate adjustment is recommended to increase precision. For important treatment effects, such as odds or hazard ratios, conditioning on covariates in binary logistic or proportional hazards models changes the interpretation of the treatment effect, and conditioning on different sets of covariates renders the resulting effect estimates incomparable. We propose a novel nonparanormal model formulation for adjusted marginal inference. This model for the joint distribution of outcome and covariates directly features a marginally defined treatment effect parameter, such as a marginal odds or hazard ratio. Not only the marginal treatment effect of interest can be estimated based on this model, it also provides an overall coefficient of determination and covariate-specific measures of prognostic strength. For the special case of Cohen's standardized mean difference d, we theoretically show that adjusting for an informative prognostic variable improves the precision of the marginal, noncollapsible effect. Empirical results confirm this not only for Cohen's d but also for odds and hazard ratios in simulations and four applications. A reference implementation is available in the R add-on package tram.
Randomized clinical trials typically aim to estimate a marginal treatment effect. While covariate adjustment can improve precision, it may change the estimand in nonlinear models due to noncollapsibility, leading to conditional rather than marginal treatment effects. At the same time, identifying prognostic and predictive covariates is important for understanding treatment effect heterogeneity and informing clinical decision-making. Keeping marginal interpretability while allowing efficiency gains and assessment of heterogeneity remains a methodological challenge. In this work, we extend nonparanormal adjusted marginal inference to allow for heterogeneous treatment effects. The proposed framework embeds the marginal treatment effect directly in a joint model for the outcome and baseline covariates. This construction preserves marginal interpretability while adjusting for potentially prognostic and/or predictive covariates. The method applies to continuous, binary, ordinal, and time-to-event outcomes and allows explicit estimation and ranking of prognostic and predictive covariates on a common scale. For continuous outcomes, we show that the asymptotic variance of the marginal treatment effect measured as Cohen's d is never worse and often better under covariate adjustment than without adjustment. Efficiency gains are primarily driven by prognostic effects, with realistic predictive effects contributing little additional improvement. Simulation studies confirm these findings across outcome types and demonstrate unbiased and more efficient estimation of marginal effects for Cohen's d, log-odds ratios, and log-hazard ratios. Application to an acupuncture trial demonstrates that the method reproduces the original trial findings while improving efficiency and allowing ranking of prognostic and predictive covariates.
Over the last five decades, we have seen strong methodological advances in survival analysis, using parametric methods and, more prominently, methods based on non-/semi-parametric estimation. As the methodological landscape continues to evolve, the task of navigating through the multitude of methods and identifying available software resources is becoming increasingly challenging-especially in more complex scenarios, such as when dealing with interval-censored or clustered survival data, non-proportional hazards, or dependent censoring. This tutorial explores the potential of using the framework of smooth transformation models for survival analysis in the R system for statistical computing. This framework provides a unified maximum-likelihood approach that covers a wide range of survival models, including well-established ones such as the Weibull model and a fully parametric version of the famous Cox proportional hazards model, and various extensions for more complex scenarios. We explore models for non-proportional/crossing hazards, dependent censoring, clustered observations and extensions towards personalized medicine within this framework. Using survival data from a two-arm randomized controlled trial on rectal cancer therapy, we demonstrate how survival analysis tasks can be seamlessly navigated in R within this framework using the implementation provided by the tram package, and few related packages.
This paper reviews and compares methods to assess treatment effect heterogeneity in the context of parametric regression models. These methods include the standard likelihood ratio tests, bootstrap likelihood ratio tests, and Goeman's global test, motivated by testing whether the random effect variance is zero. We place particular emphasis on tests based on the score-residual of the treatment effect and explore different variants of tests in this class. All approaches are compared in a simulation study, and the approach based on residual scores is illustrated in a clinical trial with a time-to-event outcome comparing treatment vs. placebo. Our findings demonstrate that score-residual-based methods provide practical, flexible, and reliable tools for exploring treatment effect heterogeneity and treatment effect modifiers, and can provide useful guidance for decision-making around treatment effect heterogeneity.
A general framework for simultaneous inference in finite mixtures of generalized linear regression models is presented. Assuming asymptotic normality of the maximum likelihood estimate of all interesting model parameters, confidence regions and p-values using a maximum norm for the multivariate t-statistic are derived. This allows to simultaneously test all regression coefficients whether they are zero. Another application is to test for constant effects across mixture components. Size and power of the new methods are evaluated using artificial data. A real world data set on the productivity of PhD students is used to demonstrate the application of the procedures.
Woody species diversity is crucial for the resilience of forests under climate change. The early stages of regeneration, particularly after canopy disturbance, shape the composition of future forest. Light availability, browsing pressure and their interactions should be key drivers of woody species diversity, biomass and density, but are not well understood due to limited experimental setups. We used exclosures in a mixed broadleaf forest with high woody diversity to protect seedlings from browsing by roe deer ( Capreolus capreolus ), the predominant ungulate in Central Europe. In a full‐factorial design, we paired 75 exclosure (fenced) and control (unfenced) plots in both closed canopy stands and experimental gaps. We measured woody regeneration 1 and 4 years after the start of the treatments. Rarefaction‐extrapolation curves revealed that diversity of common (Shannon diversity) woody species was highest in sunny and shaded exclosures (7), compared to sunny (4.5) and shaded (3.5) controls, for species that had outgrown the 130 cm browsing‐susceptible height of roe deer. Stage‐structured matrix models showed that saplings ≤20 cm had the highest probability of reaching >130 cm within 3 years in sunny exclosures (1.51%), followed by sunny controls (0.32%), shaded exclosures (0.27%) and shaded controls (0.07%). Browsing resulted in homogenized woody regeneration, particularly in forest gaps. Increasing light availability did not compensate for diversity loss by browsing. Given the densities of roe deer in the study area—typical for many Central European forests—effective population control or fencing appears essential to maintain high diversity of future forest shaping and forestry relevant woody species, especially in open forests. Synthesis and applications . Our results provide strong evidence that in the current context of increasing tree mortality and subsequent light availability, roe deer act as a keystone species. It is crucial to implement browsing protection measures and control roe deer population before or immediately after canopy disturbance. Otherwise, the rapid growth of a few dominant, browsing‐resistant woody species will significantly reduce the woody plant diversity of future forests.
Identifying reliable biomarkers for predicting clinical events in longitudinal studies is important for accurate disease prognosis and for guiding development of new treatments. However, prognostic studies are often observational, making it difficult to account for patient heterogeneity. In amyotrophic lateral sclerosis (ALS), factors such as age, site of onset and genetic status influence both survival and biomarker levels, yet their impact on the prognostic accuracy of biomarkers over time remains unclear. While time-dependent receiver operating characteristic methods have been developed to handle censored time-to-event outcomes, most do not adjust for covariates. To address this, we propose the nonparanormal prognostic biomarker framework, which models the joint distribution of the biomarker and event time while accounting for covariates. This allows estimation of covariate-specific time-dependent ROC curves and related summary measures. We apply the NPB framework to evaluate serum neurofilament light as a prognostic biomarker in ALS, showing that its accuracy varies over time and with patient characteristics. By capturing these covariate-specific effects, the NPB framework supports more targeted risk stratification and can potentially improve the design of clinical trials for new ALS treatments.
The use of metabarcoding for insect species identification has grown rapidly, but the absence of abundance data hinders meaningful diversity metrics like sample coverage-standardized species richness. Additionally, the vast number of taxa often lacks a unified phylogeny or trait database. We present a framework for constructing a phylogenetic tree encompassing the majority of insect families, standardisation of sample coverage (an objective measure of sample completeness) and assessment of both taxonomic and phylogenetic diversity using the Hill series for metabarcoding data. Applied to central Europe, our framework analysed insect diversity from 400 families along a land-use gradient. Results revealed land-use intensity significantly affects sample coverage, emphasizing the need for biodiversity standardization. After standardization, taxonomic diversity declined by 27 to 44%, and phylogenetic diversity by 13 to 29% across 39,000 Operational Taxonomic Units, from forests to agricultural areas. Rare species exhibited greater phylogenetic diversity loss than taxonomic, while dominant species showed smaller phylogenetic losses but stronger declines in taxonomic diversity. Our findings underscore agriculture's detrimental effects on specific insect taxa, even after adjusting for sample coverage, and provide new insights into the loss of functional diversity, as represented by phylogeny. ### Competing Interest Statement The authors have declared no competing interest.
Medical imaging data and electronic health records are an integral part of clinical routine and research for prognostication of patient survival and thus directly inform patient management. However, standard regression models used to derive patient prognoses are ill-equipped to handle such non-tabular data directly. Several neural network architectures based on classification or the Cox model have been proposed. Here, we present deep conditional transformation models (DCTMs) for survival applications with medical imaging data. DCTMs include the Cox model as a special case, but parameterize the log cumulative baseline hazards via Bernstein polynomials and allow the specification of nonlinear and non-proportional hazards for both tabular and non-tabular data and extend to all types of uninformative censoring. DCTMs yield moderate to large performance gains over state-of-the-art deep learning approaches to survival analysis on a multitude of publicly available datasets featuring tabular or imaging data from radiology and pathology.
Reproducibility of statistical simulations is crucial but proved being a challenge in its own right. Recently, lack of reproducibility of important simulation studies stimulated developments of reporting guidelines and specific protocols trying to improve on this situation. The problem, of course, is not new and issues regarding reproducibility of numerical results, for example statistical analyses or simulations, have long been known. Documented lack of progress regarding reproducibility in the past decade naturally leads to the question if problem awareness was not powerful enough to lead to improved reproducibility. As a benchmark case, I tried to reproduce a simulation study in a so far unpublished manuscript by Fritz Leisch and myself. The results show that, time and again, the devil is in the details and much self-discipline and extensive record-keeping and documentation are mission critical.
Treatment effects for assessing the efficacy of a novel therapy are typically defined as measures comparing the marginal outcome distributions observed in two or more study arms. Although one can estimate such effects from the observed outcome distributions obtained from proper randomisation, covariate adjustment is recommended to increase precision in randomised clinical trials. For important treatment effects, such as odds or hazard ratios, conditioning on covariates in binary logistic or proportional hazards models changes the interpretation of the treatment effect under noncollapsibility and conditioning on different sets of covariates renders the resulting effect estimates incomparable. We propose a novel nonparanormal model formulation for adjusted marginal inference. This model for the joint distribution of outcome and covariates directly features a marginally defined treatment effect parameter, such as a marginal odds or hazard ratio. Marginal distributions are modelled by transformation models allowing broad applicability to diverse outcome types. Joint maximum likelihood estimation of all model parameters is performed. From the parameters not only the marginal treatment effect of interest can be identified but also an overall coefficient of determination and covariate-specific measures of prognostic strength can be derived. A reference implementation of this novel method is available in R add-on package tram. For the special case of Cohen's standardised mean difference d, we theoretically show that adjusting for an informative prognostic variable improves the precision of this marginal, noncollapsible effect. Empirical results confirm this not only for Cohen's d but also for log-odds ratios and log-hazard ratios in simulations and three applications.