Additive models for regression functions and logistic regression functions are considered in which the component functions are fitted by cubic splines constrained to be linear in the tails. Rules and strategies for knot placement are discussed, and two illustrative applications of the resulting methodology are presented.
During the period 1962--1964, I had a tenure track Assistant Professorship in Mathematics at Cornell University in Ithaca, New York, where I did research in probability theory, especially on linear diffusion processes. Being somewhat lonely there and not liking the cold winter weather, I decided around the beginning of 1964 to try to get a job in the Mathematics Department at UCLA, in the city in which I was born and raised. At that time, Leo Breiman was an Associate Professor in that department. Presumably, he liked my research on linear diffusion processes and other research as well, since the department offered me a tenure track Assistant Professorship, which I happily accepted. During the Summer of 1965, I worked on various projects with Sidney Port, then at RAND Corporation, especially on random walks and related material. I was promoted to Associate Professor, effective in Fall, 1966, presumably thanks in part to Leo. Early in 1966, I~was surprised to be asked by Leo to participate in a department meeting called to discuss the possible hiring of Sidney. The conclusion was that Sidney was hired as Associate Professor in the department, as of Fall, 1966. Leo communicated to me his view that he thought that Sidney and I worked well together, which is why he had urged the department to hire Sidney. Anyhow, Sidney and I had a very fruitful and enjoyable collaboration in probability and, to a much lesser extent, in theoretical statistics, for a number of years thereafter.
REVIEWS 255 F. K. Kent’s work is an outstanding contribution to Renaissance studies: it is dedicated to Lorenzo De’ Medici (1449–1492), and to his patronage of the arts. Starting from a historical point of view, this book is one of the most exhaustive works dedicated to Lorenzo De’ Medici in the English language. F. K. Kent studies Lorenzo’s relationship to all the arts, and cultural circles starting from the Magnifico’s role of political boss (maestro della bottega) of republican Florence, and his fundamental position in Renaissance Italian diplomacy. Organized chronologically, each chapter of the book focuses on the different stages of Lorenzo De’ Medici’s life, as politician, businessman, and artist. Chapter 1, “ Introduction: The myth of Lorenzo,” is a brief and general introduction of the role of Lorenzo as political leader and promoter of arts. Chapter 2, “The Aesthetic Education of Lorenzo,” is dedicated to his youth, and how he developed his aesthetic, and artistic taste in relation with the other members of his family—such as his grandfather, Cosimo, and his mother, Lucrezia Tornabuoni—and the numerous contemporary artists and writers. Lorenzo became himself a gran maestro della bottega, and directed artistic production, important for his own political propaganda intended to obtain “magnificence.” Chapter 3, “The temptation to be Magnificent, 1468–1484,” always focusing on the artistic and literary productions, deals with Lorenzo’s foreign politics, the relationship with the other Italian states, and lords. Chapter 4, “Lorenzo and the Florentine Building Boom, 1485–1492,” talks about three decades of a major building program (palaces, churches, monuments) undertaken in order to have an urban renovatio based on ancient Rome. It provides very interesting material for the history of architecture, and urban planning in the Renaissance. Lorenzo read Alberti’s De re aedificatoria, and always brought discussion in his cultural circle to the topic of which architectural models to choose. Chapter 5, “Lorenzo, Fine Husbandman and Villa Builder, 1483–1492,” is devoted to Medici’s villas in the Tuscan countryside, in relation to the bucolic model from classical antiquity . Written in a very clear prose and illustrated with photographs of several artistic works commissioned by Medici, this book provides a firm historical context and precise chronology of Lorenzo de’ Medici’s relationship with art. As the author says, it “remains a historian’s contribution to art-historical debate on Lorenzo de’ Medici and the visual arts” (xi). Providing a vast bibliography, F. W. Kent’s book is also an excellent reference for further studies on the many aspects of Lorenzo’s figure, or his times, very useful to scholars of history, art history, and literature. ROSSELLA PESCATORI, Italian, UCLA The Letters of Peter Damian, 121–150, trans. Owen J. Blum and Irven M. Resnick (Washington, DC: Catholic University of America Press 2004) xxvi + 195 pp. In a letter, dated 1067, to the citizens of Florence, Peter Damian essays to repair the relations between the Florentine bishop Peter and the townsfolk. He therein asserts the authority of his written word, putting hearsay to rest and explaining, “lest my reputation be unjustly impugned … let me put in writing what I often told you in your presence, so that what you heard me say, you may now see in written form” (Letter 146). At other moments, Damian utilizes both REVIEWS 256 the conciliatory and admonitory efficacy of words on paper, comforting and advising the empress Agnes, for whom “the presence of others for conversation is wanting” (colloquentium nunc deesse praesentiam, 124), to find solace in Christ and railing against supporters of marriage within the clergy with the veiled threat, “I consider it superfluous to unsheathe the sword of my own words against you” (proprii sermonis adversum vos evaginare mucronem superfluum ducimus, 141). With such an obvious acknowledgement of the power of the written word, not to mention the correspondent’s powerful and opinionated role in the monastic life of eleventh century Europe, it should come as no surprise that the present tome of Peter Damian’s translated letters represents the fifth in a series containing his corpus of 180 epistles, edited in the original tongue for the Monumenta Germaniae Historica. Blum and Resnick thus offer not only a...
Many problems of practical interest can be formulated as the nonparametric estimation of a certain function such as a regression function, logistic or other generalized regression function, density function, conditional density function, hazard function, or conditional hazard function. Extended linear modeling provides a convenient theoretical framework for using polynomial splines and their selected tensor products in such function estimation problems and especially for obtaining rates of convergence of the resulting estimates in a unified manner. For a long time the theoretical results were restricted to fixed knot splines and to log-likelihood functions that were twice continuously differentiable. Recently, Stone and Huang extended the theory to handle free knot splines. In the present paper, the theory is further extended to handle contexts in which the log-likelihood function may not be differentiable. Specifically, we establish rates of convergence for estimation based on free knot splines in the context of nonparametric regression corresponding to M-estimates, which includes least absolute deviations (LAD) regression, quantile regression, and robust regression as special cases.
In earlier articles, we developed an automated methodology for using cubic splines with tail linear constraints to model the logarithm of a univariate density function. This methodology was subsequently modified so that the knots were determined by stepwise addition-deletion and the remaining coefficients were determined by maximum likelihood estimation. An alternative approach, referred to as the free knot spline procedure, is to use the maximum likelihood method to estimate the knot locations as well as the remaining coefficients. This article compares various approaches to constructing confidence intervals for logspline density estimates, for both the stepwise procedure and the free knot procedure. It is concluded that a variation of the bootstrap, in which only a limited number of bootstrap simulations are used to estimate standard errors that are combined with standard normal quantiles, seems to perform the best, especially when coverages and computing time are both taken into account.
In survival analysis, proportional hazards models (Cox regression) are commonly used to estimate covariate eeects. Two advantages of this approach are that the interpretation of the results is similar to that for ordinary linear models and the eeects are estimated regardless of the baseline hazard function. Inspired by the success of polynomial splines and their tensor products in adaptive multiple regression (MARS, Friedmann6]), Kooperberg, Stone and Truong 11] developed a similar adaptive hazard regression (HARE) methodology for estimating the conditional log-hazard function based on possibly censored survival data with one or more covariates. This methodology circumvents the proportionality used in proportional hazards models while still retaining the usual interpretation of the estimated eeects. HARE also provides greater exibility in modeling these eeects through the use of polynomial splines and stepwise addition and deletion of basis functions. Early attempts to use splines in survival analysis are described Other nonparametric methods, such as kernel estimates, have been used to test for nonproportionalityy7]. Intrator and Kooperbergg10] compare the use of trees and splines in survival analysis. In this entry, hazard regression refers to the HARE methodologyy11]. Some authors use hazard(s) regression or proportional hazard(s) regression to refer to the Cox proportional hazards modell3]. Let T be a (nonnegative) survival time whose distribution may depend on a vector x = (x 1 ; : : : ; x M) of covariates ranging over a subset X = X 1 X M of R M. Let f(tjx), F (tjx) = R t 0 f(ujx)du, (tjx) = f(tjx)=1 ? F (tjx)] and (tjx) = log (tjx) denote the corresponding conditional density, distribution, hazard and log-hazard functions, respectively. Let G be a p-dimensional linear space of functions on 0; 1) X, and let B 1 ; : : : ; B p be a basis of G. The HARE model for the log-hazard function is given by (tjx;) = p X j=1 j B j (tjx); t 0: (1) The coeecient vector = (1 ; : : : ; p) in (1) is estimated by the maximum likelihood method. Speciically, consider n randomly selected individuals.
Several ways to obtain pointwise confidence intervals corresponding to logspline density estimation are studied. These methods include a variety of approaches based oil estimation using free knot splines, a couple of approaches based oil the bootstrap, and a Bayesian approach. It is concluded that a variation of the bootstrap, in which only a limited number of bootstrap simulations are used to estimate standard errors that are combined with standard normal quantiles, seems to perform the best, especially when coverages and computing time are both taken into account.
Extended linear models form a very general framework for statistical modeling. Many practically important contexts fit into this framework, including regression, logistic and Poisson regression, density estimation, spectral density estimation, and conditional density estimation. Moreover, hazard regression, proportional hazard regression, marked point process regression, and diffusion processes, all perhaps with time-dependent covariates, also fit into this framework. Polynomial splines and their tensor products provide a universal tool for constructing maximum likelihood estimates for extended linear models. The theory of rates of convergence for such estimates as it applies both to fixed knot splines and to free knot splines will be surveyed, and the implications of this theory for the development of corresponding methodology will be discussed briefly.
We consider the nonparametric estimation of the drift coefficient in a diffusion type process in which the diffusion coefficient is known and the drift coefficient depends in an unknown manner on a vector of time-dependent covariates. Based on many continuous realizations of the process, the estimator is constructed using the method of maximum likelihood, where the maximization is taken over a finite dimensional estimation space whose dimension grows with the sample size n. We focus on estimation spaces of polynomial splines. We obtain rates of convergence of the spline estimates when the knot positions are prespecified but the number of knots increases with the sample size. We also give the rates of convergence for free knot spline estimates, in which the knot positions of splines are treated as free parameters that are determined by data.
Many problems of practical interest can be formulated as the estimation of a certain function such as a regression function, logistic or other generalized regression function, density function, conditional density function, hazard function, or conditional hazard function. Extended linear modeling provides a convenient framework for using polynomial splines and their tensor products in such function estimation problems. Huang (Statist. Sinica 11 (2001) 173) has given a general treatment of the rates of convergence of maximum likelihood estimation in the context of concave extended linear modeling. Here these results are generalized to let the approximation space used in the fitting procedure depend on a vector of parameters. More detailed treatments are given for density estimation and generalized regression (including ordinary regression) on the one hand and for approximation spaces whose components are suitably regular free knot splines and their tensor products on the other hand.
The logarithm of the relative risk function in a proportional hazards model involving one or more possibly time-dependent covariates is treated as a specified sum of a constant term, main effects, and selected interaction terms. Maximum partial likelihood estimation is used, where the maximization is taken over a suitably chosen finite-dimensional estimation space, whose dimension increases with the sample size and which is constructed from linear spaces of functions of one covariate and their tensor products. The L-2 rate of convergence for the estimate and its ANOVA components is obtained. An adaptive numerical implementation is discussed, whose performance is compared to (full likelihood) hazard regression both with and without the restriction to proportional hazards.
Kooperberg, Bose, and Stone introduced POLYCLASS, a methodology that uses adaptively selected linear splines and their tensor products to model conditional class probabilities. The authors attempted to develop a methodology that would work well on small and moderate size problems and would scale up to large problems. However, the version of POLYCLASS that was developed for large problems was computationally impractical beyond a certain point. In this article we gain further insight into the fitting of large POLYCLASS models by simultaneously considering the fitting of large feed-forward neural network models with a single hidden layer (NEURALNET). In this combined setting, the stochastic gradient method, as used in the online version of the backpropagation method for fitting neural network models, and a stochastic version of the conjugate gradient method for fitting such models emerge as being computationally attractive. In particular, these stochastic methods are successfully applied to the fitting of POLYCLASS and NEURALNET models in the context of a phoneme recognition problem involving 45 phonemes, 81 features, 150,000 cases in the training sample, up to 1,000 basis functions and 44,000 parameters for POLYCLASS, and up to 800 hidden nodes and about 100,000 parameters for NEURALNET.
Consider repeated events of multiple kinds that occur according to a right‐continuous semi‐Markov process whose transition rates are influenced by one or more time‐dependent covariates. The logarithms of the intensities of the transitions from one state to another are modelled as members of a linear function space, which may be finite‐ or infinite‐dimensional. Maximum likelihood estimates are used, where the maximizations are taken over suitably chosen finite‐dimensional approximating spaces. It is shown that the L2 rates of convergence of the maximum likelihood estimates are determined by the approximation power and dimension of the approximating spaces. The theory is applied to a functional ANOVA model, where the logarithms of the intensities are approximated by functions having the form of a specified sum of a constant term, main effects (functions of one variable), and interaction terms (functions of two or more variables). It is shown that the curse of dimensionality can be ameliorated if only main effects and low‐order interactions are considered in functional ANOVA models.