Our model for the lifespan of an enterprise is the geometric distribution. We do not formulate a model for enterprise foundation, but assume that foundations and lifespans are independent. We aim to fit the model to information about foundation and closure of enterprises in a panel of the German statistical offices. The lifespan for an enterprise that has been founded before the first wave of the panel is either observable, when the enterprise is contained in the panel, or missing, when it already closed down before the first wave. Marginalizing the likelihood to that part of the enterprise history after the first wave contributes to the aim of closed-form estimate and standard error. Invariance under the foundation distribution is achieved by conditioning on observability of the enterprises. The conditional marginal likelihood can be written as a function of a martingale. The later arises when calculating the compensator, with respect some filtration, of a process that counts the closures. The estimator itself can then also be written as a martingale transform and consistency as well as asymptotic normality are easily proven. The life expectancy of German enterprises, estimated from the demographic information about 1.4 million enterprises for the years 2018 and 2019, are ten years. The width of the confidence interval are three to four weeks. Closure after the last wave is taken into account as right censored.
Given a double-truncated sample of lifespans, we test the hypothesis of a parametric distribution family for the lifespan. Demography typically certifies the life expectancy to be nonstationary in time. We model the resulting dependence between a lifespan and the birthday of an individual with a copula. Our main example is the Farlie-Gumbel-Morgenstern copula. The asymptotic null distribution of the test is based on Donsker-class arguments and the functional delta method for empirical processes. One assumption for the test is consistency in the estimation of the parameter for the model under the null hypothesis. Requirements for consistency are fulfilled by our set of assumptions. The test is carried out in two stages, the computation of the test statistics and the simulation of the critical value. The statistic can be reached after finitely many computation steps, whereas the critical value can only be an approximation. With the exponential distribution as an example for the lifespan distribution, and for the application to 55,000 German double-truncated enterprise lifespans, the constructed Kolmogorov-Smirnov test rejects clearly an age-homogeneous closure hazard.
In studies on lifetimes, occasionally, the population contains statistical units that are born before the data collection has started. Left-truncated are units that deceased before this start. For all other units, the age at the study start often is recorded and we aim at testing whether this second measurement is independent of the genuine measure of interest, the lifetime. Our basic model of dependence is the one-parameter Gumbel-Barnett copula. For simplicity, the marginal distribution of the lifetime is assumed to be Exponential and for the age-at-study-start, namely the distribution of birth dates, we assume a Uniform. Also for simplicity, and to fit our application, we assume that units that die later than our study period, are also truncated. As a result from point process theory, we can approximate the truncated sample by a Poisson process and thereby derive its likelihood. Identification, consistency and asymptotic distribution of the maximum-likelihood estimator are derived. Testing for positive truncation dependence must include the hypothetical independence which coincides with the boundary of the copula’s parameter space. By non-standard theory, the maximum likelihood estimator of the exponential and the copula parameter is distributed as a mixture of a two- and a one-dimensional normal distribution. For the proof, the third parameter, the unobservable sample size, is profiled out. An interesting result is, that it differs to view the data as truncated sample, or, as simple sample from the truncated population, but not by much. The application are 55 thousand double-truncated lifetimes of German businesses that closed down over the period 2014 to 2016. The likelihood has its maximum for the copula parameter at the parameter space boundary so that the p-value of test is 0.5. The life expectancy does not increase relative to the year of foundation. Using a Farlie–Gumbel–Morgenstern copula, which models positive and negative dependence, finds that life expectancy of German enterprises even decreases significantly over time. A simulation under the condition of the application suggests that the tests retain the nominal level and have good power.
We approximate a small truncated sample of censored durations with a bivariate Poisson process. For exponential or geometric distribution of the duration, we derive the likelihood, profile out the unobservable sample size, and study identification, consistency as well as asymptotic normality.
With double-truncated lifespans, we test the hypothesis of a parametric distribution family for the lifespan. The typical finding from demography is an instationary behaviour of the life expectancy, and a copula models the resulting weak dependence of lifespan and the age at truncation. Our main example is the Farlie-Gumbel-Morgenststern copula. The test is based on Donsker-class arguments and the functional delta method for empirical processes. The assumptions also allow parametric inference, and proofs slightly simplify due to the compact support of the observations. An algorithm with finitely many operations is given for the computation of the test statistic. Simulations becomes necessary for computing the critical value. With the exponential distribution as an example, and for the application to 55{,}000 German double-truncated enterprise lifespans, the constructed Kolmogorov-Smirnov test rejects clearly an age-homogeneous closure hazard.
A continuous-time multi-state history is semi-Markovian, if an intensity to migrate from one state into another, depends on the duration in the first state. Such duration can be formalised as covariate, entering the intensity process of the transition counts. We derive the integrated intensity process, prove its predictability and the martingale property of the residual. In particular, we verify the usual conditions for the respective filtration. As a consequence, according to Nielsen and Linton (1995), a kernel estimator of the transition intensity, including the duration dependence, converges point-wise at a slow rate, compared to the Markovian kernel estimator, i.e when ignoring dependence. By using the rate discrepancy, we follow Gozalo (1993) and show that the (properly scaled) maximal difference of the two kernel estimators on a random grid of points is asymptotically chi-square-1-distributed. As a data example, for a sample of 130,000 German women observed over a period of nine years, we model the mortality after dementia onset, potentially dependent on the disease duration. As usual, the models under both hypotheses need to be enlarged to allow for independent right-censoring. We find a significant effect of dementia duration, nearly independent of the bandwidth.
From the inventory of the health insurer AOK in 2004, we draw a sample of a quarter million people and follow each person’s health claims continuously until 2013. Our aim is to estimate the effect of a stroke on the dementia onset probability for Germans born in the first half of the 20th century. People deceased before 2004 are randomly left-truncated, and especially their number is unknown. Filtrations, modelling the missing data, enable circumventing the unknown number of truncated persons by using a conditional likelihood. Dementia onset after 2013 is a fixed right-censoring event. For each observed health history, Jacod’s formula yields its conditional likelihood contribution. Asymptotic normality of the estimated intensities is derived, related to a sample size definition including the number of truncated people. The standard error results from the asymptotic normality and is easily computable, despite the unknown sample size. The claims data reveal that after a stroke, with time measured in years, the intensity of dementia onset increases from 0.02 to 0.07. Using the independence of the two estimated intensities, a 95% confidence interval for their difference is [0.053, 0.057]. The effect halves when we extend the analysis to an age-inhomogeneous model, but does not change further when we additionally adjust for multi-morbidity.
When estimating a proportion and only a sample of triplets is given, dependencies within the triplets are to be accounted for. Without assuming a distribution for the success count of the triplet, together with the proportion, as second and third parameter the correlations of 1st and 2nd order enter the model. We apply maximum likelihood estimation, and derive consistency by using that the triplet count is multinomially distributed, combined with the continuous mapping theorem. The asymptotic normality follows with the delta-method, resulting in closed-form expressions for the standard errors. As application we study caries prevalence of pre-school children from a sample to nursing schools. We compare the standard errors with those for assuming erroneously independence within the nursing schools. As to be suspected, the design `inflates' the standard error markedly.
Truncated survival data are observed retrospectively, if the death event falls into the study period. We find that estimating the survival distribution requires a model for the birth event. With births generated by a parametric Poisson process, we analyse the likelihood of parametric survival, stochastically independent of birth. Conditioning on the number of observed units reduces the number of parameters by one, and to two in our application. The compact support of an observation simplifies the proof of consistency. Furthermore, that we only need to show identification separately for the birth and the death distribution is helpful for demonstrating asymptotic normality. For identification of the survival parameters, a stronger criterion is needed than for a simple random sample, but is fulfilled in our application. In a simulation study, we find that the variance inflation by truncation can be substantial, and apparently is indeed so in our application. From 55,000 German companies that went insolvent between 2014 and 2016, we infer an average time to insolvency of six years and a negative linear trend of corporate foundations after the German reunification in 1990. Companies in Northern Germany survive longer than in the south.
For a sample of Exponentially distributed durations we aim at point estimation and a confidence interval for its parameter. A duration is only observed if it has ended within a certain time interval, determined by a Uniform distribution. Hence, the data is a truncated empirical process that we can approximate by a Poisson process when only a small portion of the sample is observed, as is the case for our applications. We derive the likelihood from standard arguments for point processes, acknowledging the size of the latent sample as the second parameter, and derive the maximum likelihood estimator for both. Consistency and asymptotic normality of the estimator for the Exponential parameter are derived from standard results on M-estimation. We compare the design with a simple random sample assumption for the observed durations. Theoretically, the derivative of the log-likelihood is less steep in the truncation-design for small parameter values, indicating a larger computational effort for root finding and a larger standard error. In applications from the social and economic sciences and in simulations, we indeed, find a moderately increased standard error when acknowledging truncation.
We observe a quarter million people over a period of nine years and are interested in the effect of a stroke on the probability of dementia onset. Randomly lefttruncated has a person been that was already deceased before the period. The ages at a stroke event or dementia onset are conditionally fixed right-censored, when either event may still occur, but after the period. We incorporate death and model the history of the three events by a homogeneous Markov process. The compensator for the respective counting processes is derived and Jacod’s formula yields the likelihood contribution, conditional on observation. An Appendix is devoted to the role of filtrations in deriving the likelihood, for the simplification of merged non-dead states. Asymptotic normality of the estimated intensities is derived by martingale theory, relative to the size of the sample including the truncated persons. The data of a German health insurance reveals that after a stroke, the intensity of dementia onset is increased from 0.02 to 0.07, for Germans born in the first half on the 20 century. The intensity difference has a 95%-confidence interval of [0.048, 0.051] and the difference halves when we extent to an age-inhomogeneous model due to Simpson’s paradox.
We model an overdispersed count as a dependent measurement, by means of the Negative Binomial distribution. We consider a quantitative covariate that is fixed by design. The expectation of the dependent variable is assumed to be a known function of a linear combination involving the possibly multidimensional covariate and its coefficients. In the NB1-parametrization of the Negative Binomial distribution, the variance is a linear function of the expectation, inflated by the dispersion parameter, and the distribution not a generalized linear model. For the maximum likelihood estimator for all parameters we apply a general result of Bradley and Gart (Biometrika 49:205–214, 1962) to derive weak consistency and asymptotic normality and a technique in Fahrmeir and Kaufmann (Ann Stat 13:342–368, 1985) for strong consistency. To this end, we show (1) how to bound the logarithmic density by a function that is linear in the outcome of the dependent variable, independently of the parameter. Furthermore (2) the positive definiteness of the matrix related to the Fisher information is shown with the Cauchy–Schwarz inequality.
We estimate the dementia incidence hazard in Germany for the birth cohorts 1900 until 1954 from a simple sample of Germany’s largest health insurance company. Followed from 2004 to 2012, 36,000 uncensored dementia incidences are observed and further 200,000 right-censored insurants included. From a multiplicative hazard model we find a positive and linear trend in the dementia hazard over the cohorts. The main focus of the study is on 11,000 left-censored persons who have already suffered from the disease in 2004. After including the left-censored observations, the slope of the trend declines markedly due to Simpson’s paradox, left-censored persons are imbalanced between the cohorts. When including left-censoring, the dementia hazard increases differently for different ages, we consider omitted covariates to be the reason. For the standard errors from large sample theory, left-censoring requires an adjustment to the conditional information matrix equality.
There is substantial evidence that bank rating data display non-Markovian structures. We introduce a non-Markovian parameter in a simple model for rating transition histories. Accounting for the frequent statistical obstacle of partially missing transitions, we make use of the expectation maximisation (EM) algorithm to estimate all model parameters and find a marked non-Markovian effect for our data.
Dieses Buch verzahnt die wesentlichen Grundlagen aus Mathematik, Wirtschaft und Statistik, die zum Verständnis von Finanzmärkten nötig sind. Die Inhalte bilden eine solide Basis, mit der sich sowohl Praktiker als auch Theoretiker weiterführende Literatur zügig selbst erschließen können.
Prof. Dr. Josef Bleymüller war Direktor des Instituts für Ökonometrie und Wirtschaftsstatistik der Universität Münster.