Discrete-time hazard models are widely used when event times are measured in intervals or are not precisely observed. While these models can be estimated using standard generalized linear model techniques, they rely on extensive data augmentation, making estimation computationally demanding in large-scale, high-dimensional settings. In this article, we demonstrate how the recently proposed batchwise backfitting algorithm, a general framework for scalable estimation and variable selection in distributional regression, can be effectively extended to discrete hazard models. Using both simulated data and a large-scale application on infant mortality in sub-Saharan Africa, we show that the algorithm delivers accurate estimates, automatically selects relevant predictors and scales efficiently to large datasets. The findings underscore the algorithm's practical utility for analyzing large-scale, complex survival data with high-dimensional covariates.
Dieses Arbeitsbuch ergänzt perfekt das Lehrbuch Fahrmeir/Künstler/Pigeot/Tutz: Statistik - Der Weg zur Datenanalyse. Es enthält die Lösungen zu den dort gestellten Aufgaben. Darüber hinaus bietet es e
Modeling real estate prices in the context of hedonic models often involves fitting a Generalized Additive Model, where only the mean of a (lognormal) distribution is regressed on a set of variables without taking other parameters of the distribution into account. Thus far, the application of regression models that model the full conditional distribution of the prices, has been infeasible for large data sets, even on powerful machines. Moreover, accounting for heterogeneity of effects regarding time and locale, is often achieved by naive stratification of the data rather than on a model basis. A novel batchwise backfitting algorithm is applied in the context of a structured additive distributional regression model, which enables us to efficiently model all distributional parameters of the price distribution. Using a large German dataset of apartment asking prices with over one million observations, we employ a model-based clustering algorithm to capture the heterogeneity of covariate effects on the parameters with respect to dwelling locale. We thus identify clusters that are homogeneous with respect to the influence of dwelling locale on price. A boosting type algorithm of the batchwise backfitting algorithm is then used to automatically determine the variables relevant for modelling the location and scale parameters in each regional cluster. This allows for a different influence of variables on the distribution of prices depending on the locale and price segment of the dwelling.
Recently, fitting probabilistic models have gained importance in many areas but estimation of such distributional models with very large data sets is a difficult task. In particular, the use of rather complex models can easily lead to memory-related efficiency problems that can make estimation infeasible even on high-performance computers. We therefore propose a novel backfitting algorithm, which is based on the ideas of stochastic gradient descent and can deal virtually with any amount of data on a conventional laptop. The algorithm performs automatic selection of variables and smoothing parameters, and its performance is in most cases superior or at least equivalent to other implementations for structured additive distributional regression, e.g., gradient boosting, while maintaining low computation time. Performance is evaluated using an extensive simulation study and an exceptionally challenging and unique example of lightning count prediction over Austria. A very large dataset with over 9 million observations and 80 covariates is used, so that a prediction model cannot be estimated with standard distributional regression methods but with our new approach.
BackgroundThe surgical treatment of insular lesions has been historically associated with high morbidity. Laser interstitial thermal therapy (LITT) has been increasingly used in the treatment of insular lesions, commonly neoplastic or epileptogenic. Stereotaxis is used to guide laser probes to the insula where real-time magnetic resonance thermometry defines lesion creation. There is an absence of previously published reviews on insular LITT, despite a rapid uptake in use, making further study imperative. MethodsHere we present a systematic review of the PubMed and Scopus databases, examining the reported clinical indications, outcomes, and adverse effects of insular LITT. ResultsA review of the literature revealed 10 retrospective studies reporting on 53 patients (43 pediatric and 10 adults) that were treated with insular LITT. 87% of cases were for the treatment of epilepsy, with 89% of patients achieving seizure outcomes of Engle I-III following treatment. The other 13% of cases reported on insular tumors and radiological improvement was seen in all cases following treatment. All but one study reported adverse events following LITT with a rate of 37%. The most common adverse events were transient hemiparesis (29%) and transient aphasia (6%). One patient experienced an intracerebral hemorrhage, which required a decompressive hemicraniectomy, with subsequent full recovery. ConclusionThis systematic review highlights the suitability of LITT for the treatment of both insular seizure foci and insular tumors. Despite the growing use of this technique, prospective studies remain absent in the literature. Future work should directly evaluate the efficacy of LITT with randomized and controlled trials.
Joint models for longitudinal and time-to-event data simultaneously model longitudinal and time-to-event information to avoid bias by combining usually a linear mixed model with a proportional hazards model. This model class has seen many developments in recent years, yet joint models including a spatial predictor are still rare and the traditional proportional hazards formulation of the time-to-event part of the model is accompanied by computational challenges. We propose a joint model with a piecewise exponential formulation of the hazard using the counting process representation of a hazard and structured additive predictors able to estimate (non-)linear, spatial and random effects. Its capabilities are assessed in a simulation study comparing our approach to an established one and highlighted by an example on physical functioning after cardiovascular events from the German Ageing Survey. The Structured Piecewise Additive Joint Model yielded good estimation performance, also and especially in spatial effects, while being double as fast as the chosen benchmark approach and performing stable in an imbalanced data setting with few events.
Real estate valuation is typically based on hedonic regression models where the expected price of a property is explained in dependence of its attributes. However, investors in the housing market are equally interested in the distribution of real estate market values (including price variation), that is, determining the impact of attributes of a property on the entire conditional distribution. We therefore consider Bayesian structured additive distributional and quantile regression models for real estate valuation. In the first approach, each parameter of a potentially complex parametric response distribution is related to a structured additive predictor. In contrast, the second approach proceeds differently and models arbitrary quantiles of the response distribution directly and nonparametrically. Both models presented are based on a multilevel version of structured additive regression thereby utilizing the typical hierarchical structure of real estate data. We demonstrate the proposed methodology within a detailed case study based on more than 3 000 owner-occupied single family homes in Austria, discuss interpretation of the resulting effect estimates, and compare models based on their predictive ability.
Modeling real estate prices in the context of hedonic models typically involves fitting a Generalized Additive Model, where only the mean of a (lognormal) distribution is regressed on a set of variables, without taking into account other parameters of the distribution. Thus far, the application of regression models that model the full conditional distribution of the prices, has been infeasible for large data sets, even on powerful machines. Moreover, accounting for heterogeneity of effects regarding time and location, is often achieved by naive stratification of the data rather than on a model basis. We apply a novel batchwise backfitting algorithm in the context of a structured additive regression model that enables us to efficiently model all distributional parameters of an appropriate distribution. Using a large German dataset of rental prices comprising over a million observations, we choose variables relevant for modeling the location and scale parameters using a boosting variant of the algorithm. Moreover, we identify heterogeneity of covariates’ effects on the parameters with respect to both time and location on a model basis. In this way, we allow varying influence of variables on the prices’ distribution depending on the dwelling’s location and the date of sale. Modeling the full distribution of prices further enables us to investigate the influence of the variables not only on the median, but also on other quantiles of rental prices.
The most widely used approaches in hedonic price modelling of real estate data and price index construction are Time Dummy and Imputation methods. Both methods, however, reveal extreme approaches regarding regression modeling of real estate data. In the time dummy approach, the data are pooled and the dependence on time is solely modelled via a (nonlinear) time effect through dummies. Possible heterogeneity of effects across time, i.e. interactions with time, are completely ignored. Hence, the approach is prone to biased estimates due to underfitting. The other extreme poses the imputation method where separate regression models are estimated for each time period. Whereas the approach naturally includes interactions with time, the method tends to overfit and therefore increased variability of estimates. In this paper, we therefore propose a generalized approach such that time dummy and imputation methods are special cases. This is achieved by reexpressing the separate regression models in the imputation method as an equivalent global regression model with interactions of all available regressors with time. Our approach is applied to a large dataseton offer prices for private single as well as semi-detached houses in Germany. More specifically, we a) compute a Time Dummy Method index based on a Generalized Additive Model allowing for smooth effects of the metric covariates on the price utilizing the pooled data set, b) construct an Imputation Approach model, where we fit a regression model separately for each time period, c) finally develop a global model that captures only relevant interactions of the covariates with time. An important methodolical aspect in developing the global model is the usage of model-based recursive partitioning trees to define data driven and parsimonious time intervals.
We establish Bayesian effect selection for the broad class of structured additive distributional regression models using a spike and slab prior specification with scaled beta prime marginals for the importance parameters of blocks of regression coefficients. This enables us to model and select effects in all distributional parameters, such as location, scale, skewness or correlation parameters, for arbitrary distributions. The regression specifications encompass various effect types such as non-linear or spatial effects. Our spike and slab prior relies on a parameter expansion that separates blocks of regression coefficients into overall scalar importance parameters and vectors of standardised coefficients, and yields effective shrinkage and good sampling performance. Using constrained priors, it is possible to implement effect decompositions, where, for example, a non-linear effect can be decomposed into a linear component and the non-linear deviation from this linear effect; and to select both separately. We investigate some shrinkage properties, propose a way of eliciting prior hyperparameters and provide full posterior inference through Markov Chain Monte Carlo simulations. Using both simulated and real data sets, we show that our approach is applicable for data with various functional covariate effects, multilevel predictors and non-standard response distributions, such as bivariate Gaussian or zero-inflated Poisson.