
The stochastic gradient descent (SGD) algorithm has been widely used to optimize deep Cox neural network (Cox-NN) by updating model parameters using mini-batches of data. We show that SGD aims to optimize the average of mini-batch partial-likelihood, which is different from the standard partial-likelihood. This distinction requires developing new statistical properties for the global optimizer, namely, the mini-batch maximum partial-likelihood estimator (mb-MPLE). We establish that mb-MPLE for Cox-NN is consistent and achieves the optimal minimax convergence rate up to a polylogarithmic factor. For Cox regression with linear covariate effects, we further show that mb-MPLE is n -consistent and asymptotically normal with asymptotic variance approaching the information lower bound as batch size increases, which is confirmed by simulation studies. Additionally, we offer practical guidance on using SGD, supported by theoretical analysis and numerical evidence. For Cox-NN, we demonstrate that the ratio of the learning rate to the batch size is critical in SGD dynamics, offering insight into hyperparameter tuning. For Cox regression, we characterize the iterative convergence of SGD, ensuring that the global optimizer, mb-MPLE, can be approximated with sufficiently many iterations. Finally, we demonstrate the effectiveness of mb-MPLE in a large-scale real-world application where the standard MPLE is intractable.
In the general setting of independent data with possibly very different distributions, extreme value estimators inevitably target the tail of the average distribution function. We consider all possible cases, that is, the extreme value index of the average distribution can be negative, zero, or positive, and we present novel asymptotic theory for the moment estimator. Our results require a different and much more challenging proof than those for the power-law case and are based on a uniform central limit theorem for the underlying weighted tail empirical process. We find that, due to the heterogeneity of the data, the asymptotic variance of the moment estimator can be much smaller than that in the iid case. We also unravel the improved performance of high quantile and endpoint estimators in this setup. In case of a heavy tail, we ameliorate the Hill estimator by taking an optimal combination of the Hill and the moment estimator. Simulations show the good finite-sample behavior of our limit results. Finally we present applications to the maximum lifespan of monozygotic twins, the ultimate 200m running world records, and to the tail heaviness of energies of earthquakes around the globe. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.