This note presents an alternative to multiple imputation and other approaches to regression analysis in the presence of missing covariate data. Our recommendation, based on factorial and fractional factorial arrangements, is more faithful to ancillarity considerations of regression analysis and involves assessing the sensitivity of inference on each regression parameter to missingness in each of the explanatory variables. The ideas are illustrated on a medical example concerned with the success of hematopoietic stem cell transplantation in children, and on a sociological example concerned with socio-economic inequalities in educational attainment.
The simplest form of retrospective study allows the reconstruction of the dependence between a binary outcome, Y , representing the contrast between cases and controls, and one or more explanatory variables. A different objective for such situations is considered, in which there are distinct explanatory variables, say ( W , X ) determining Y . Reconstruction of the originating distribution of ( W , X ) from the case-control data is considered for both continuous and binary variables. Emphasis is on the linear regression coefficient of W on X . That coefficient, but not the relevant intercept, shows considerable stability, as shown by theory and simulations. An approximation to the value of the coefficient not conditioning on Y is given.
Models whose associated likelihood functions fruitfully factorise are an important minority allowing elimination of nuisance parameters via partial likelihood, an operation that is valuable in both Bayesian and frequentist inferences, particularly when the number of nuisance parameters is not small. After some general discussion of partial likelihood, we focus on marginal likelihood factorisations, which are particularly difficult to ascertain from elementary calculations. We suggest a systematic approach for deducing transformations of the data, if they exist, whose marginal likelihood functions are free of the nuisance parameters. This is based on the solution to an integro-differential equation constructed from aspects of the Laplace transform of the probability density function, for which candidate solutions solve a simpler first-order linear homogeneous differential equation. The approach is generalised to the situation in which such factorisable structure is not exactly present. Examples are used in illustration. Although motivated by inferential problems in statistics, the proposed construction is of independent interest and may find application elsewhere.
Background Airway inflammation promotes bronchiectasis and lung injury in cystic fibrosis (CF). Amplification of inflammation underlies pulmonary exacerbations of disease. We asked whether sputum inflammatory biomarkers provide explanatory information on pulmonary exacerbations. Patients and Methods We collected sputum from randomly chosen stable adolescents and adults and prospectively observed time to next exacerbation, our primary outcome. We evaluated relationships between potential biomarkers of inflammation, clinical characteristics and outcomes and assessed clinical variables as potential confounders or mediators of explanatory models. We assessed associations between the markers and time to next exacerbation using proportional hazard models adjusting for confounders. Results We enrolled 114 patients, collected data on clinical variables [December 8, 2014 to January 16, 2016; 46% male, mean age 28 years (SD 12), mean percent predicted forced expiratory volume in 1 s (FEV1%) 70 (SD 22)] and measured 24 inflammatory markers. Half of the inflammatory markers were plausibly associated with time to next exacerbation. Age and sex were confounders while we found that FEV1% was a mediator. Three potential biomarkers of RAGE axis inflammation were associated with time to next exacerbation while six potential neutrophil-associated biomarkers indicate associations between protease activity or reactive oxygen species with time to next exacerbation. Conclusion Pulmonary exacerbation biomarkers are part of the RAGE proinflammatory axis or reflect neutrophil activity, specifically implicating protease and oxidative stress injury. Further investigations or development of novel anti-inflammatory agents should consider RAGE axis, protease and oxidant stress antagonists. Tweetable abstract Sputum from 114 randomly chosen people with CF show RAGE axis inflammation, protease and oxidative stress injury are associated with time to next pulmonary exacerbation and may be targets for bench or factorial design interventional studies. (242 characters) ### Competing Interest Statement TGL, JAF, JLJ, YL, KAP and JBV received other support from the CFF (CC132-16AD, LIOU14Y0, LIOU14P0) and the National Heart Lung and Blood Institute (NHLBI) of the National Institutes of Health (NIH) (R01 HL125520) and received support during the current study for performing clinical trials from Abbvie, Calithera Biosciences, Corbus Pharmaceuticals, Gilead Sciences, Laurent Pharmaceuticals, Nivalis Therapeutics, Novartis, Proteostasis, Savara Pharmaceuticals, Translate Bio and Vertex Pharmaceuticals. FRA received additional other support from the NHLBI/NIH (R01 HL125520), the National Science Foundation (EMSW21-RTG) and the Margolis Foundation of Utah. PSB received other support from the CFF (Center and TDC grants) and the NHLBI/NIH (U01 HL114623) and received support for a clinical trial from Alcresta Therapeutics. BAC received other support from the CFF (C112-12, C112-TDC09Y, 10063SUB, 41339154.s132P010379SUB) and received support for clinical trials from Genentech, Novartis and Vertex Pharmaceuticals. CLD received other support from the CFF (C004-11, C004-TDC09Y, DAINES11Y3) and from the Health Resources and Services Administration (T72MC00012). JAF is now an employee of ICON plc, a clinical research organization involved in various trials pertinent to CF; She and ICON had no direct involvement in performance of the study following the change in affiliation. TH received other support from the CFF (PACE, Center Grant) and received support for clinical trials from Celtaxsys and Vertex Pharmaceuticals. JRH received other support from the NHLBI/NIH (HHSN268200900018C) and the Veterans Administration Healthcare System (I01 BX001533). JL received other support from the CFF (C017-11AF). CN received other support from the CFF (C138-12). PR received other support from the CFF (C003-12, C003-TDC09Y). SDS received other support from the CFF (AQUADEK12K1, SAGEL11CS0, GOAL13K2, NICK13A0, SAGEL14K1, NICK15R0) and the NHLBI/NIH (U54 HL096458) and the NCATS/NIH (Colorado CTSA Grant Number UL1 TR002535). JLT-C received other support from the CFF (TDC) and the NHLBI/NIH (HL103801) and received support for clinical trials from Vertex Pharmaceuticals. KNO is funded by the intramural research program of the NHLBI, NIH. ### Clinical Protocols ### Funding Statement This project was supported by the CF Foundation (CFF) (LIOU13A0, LIOU14Y4), the National Center for Advancing Translational Science at the National Institutes of Health (NCATS/NIH 8UL1TR000105 [formerly UL1RR025764]), the Ben B and Iris M Margolis Foundation of Utah and the Claudia Ruth Goodrich Stevens Endowment Fund. Neither the project sponsors nor any sources of other support had direct roles in development and conduct of the study. None of the sponsors of clinical trials mentioned in the competing interests and other support section that follows participated in any way with this trial. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The IRB of the University of Utah gave ethical approval for this work. The Institutional Review Board of St Luke's Health System gave ethical approval for this work. The Western Institutional Review Board for the Las Vegas CF Center gave ethical approval for this work. The Institutional Review Board of National Jewish Health gave ethical approval for this work. The Colorado Multiple Institutional Review Board for Children's Hospital Colorado gave ethical approval for this work. The Institutional Review Board of Phoenix Children's Hospital gave ethical approval for this work. The IRB of the University of Arizona gave ethical approval for this work. The Human Research Review Committee in the Human Research Protections Office of the University of New Mexico gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data produced in the present study will be available upon reasonable request to the authors after peer-reviewed publication.
With very large amounts of data, important aspects of statistical analysis may appear largely descriptive in that the role of probability sometimes seems limited or totally absent. The main emphasis of the present paper lies on contexts where formulation in terms of a probabilistic model is feasible and fruitful but to be at all realistic large numbers of unknown parameters need consideration. Then many of the standard approaches to statistical analysis, for instance direct application of the method of maximum likelihood, or the use of flat priors, often encounter difficulties. After a brief discussion of broad conceptual issues, we provide some new perspectives on aspects of high-dimensional statistical theory, emphasizing a number of open problems.
To examine innate immune responses in early SARS-CoV-2 infection that may change clinical outcomes, we compared nasopharyngeal swab data from 20 virus-positive and 20 virus-negative individuals. Multiple innate immune-related and ACE-2 transcripts increased with infection and were strongly associated with increasing viral load. We found widespread discrepancies between transcription and translation. Interferon proteins were unchanged or decreased in infected samples suggesting virally-induced shut-off of host anti-viral protein responses. However, IP-10 and several interferon-stimulated gene proteins increased with viral load. Older age was associated with modifications of some effects. Our findings may characterize the disrupted immune landscape of early disease.
A broad review is given of some areas of multivariate analysis that are not frequently emphasized. We start with situations in which underlying distributions are far from multivariate normal form, so that standard methods of multivariate analysis based on covariances are likely to be unsatisfactory. We emphasize the important distinction between internal and external analyses associated with multiple outcomes. A second broad theme relates to multiple outcomes generated by time or spatial series in which long-range dependence operates. Some implications are summarized.
Parametric statistical problems involving both large amounts of data and models with many parameters raise issues that are explicitly or implicitly differential geometric. When the number of nuisance parameters is comparable to the sample size, alternative approaches to inference on interest parameters treat the nuisance parameters either as random variables or as arbitrary constants. The two approaches are compared in the context of parametric survival analysis, with emphasis on the effects of misspecification of the random effects distribution. Notably, we derive a detailed expression for the precision of the maximum likelihood estimator of an interest parameter when the assumed random effects model is erroneous, recovering simply derived results based on the Fisher information in the correctly specified situation but otherwise illustrating complex dependence on other aspects. Methods of assessing model adequacy are given. The results are both directly applicable and illustrate general principles of inference when there is a high-dimensional nuisance parameter. Open problems with an information geometrical bearing are outlined.
It is a privilege to have the chance of congratulating Professor Efron first on this wise paper, then on the richly merited International Prize, the award of which the paper commemorates, but above all on the whole body of his deeply impressive, wide-ranging contributions to our subject. One issue which the paper indirectly raises is the role of different approaches to the conceptual and mathematical theory of our field. One of the appeals of the field is its totally international character yet, inevitably, broad and hazily defined national contrasts are visible. Thus Professor Efron links “traditional” statistical theory to Neyman and Pearson. Yet the two set out to clarify earlier work of R.A. Fisher, first with Fisher’s encouragement, which only later turned to destructive hostility. The mathematical clarity of Neyman’s work is, of course, appealing but it may be argued that its overformalization continues to lead to misunderstanding, unproductive discussion and rigidity concerning, in particular, the role of significance tests. The Nordic approaches have made an important and distinctive contribution to the field. See, in particular, the recent fine account of Rolf Sundberg, Statistical modelling by exponential families, stemming from the much earlier contribution of Per Martin-Lof. In the UK, Egon Pearson played a distinctive role in the development of our subject. He mostly had a preference for numerical illustration and was deeply involved in applications, stemming in part from an early visit he made to Bell Labs. One of the great masterpieces of our field, largely unread nowadays, is M.S. Bartlett’s (1958) Introduction to stochastic processes and their application and Bartlett’s influence on statistical development, at least in UK, was second only to Fisher’s. See, for example, his treatment in 1937 of asymptotic theory in which the dimension of the parameter space increases proportionally to sample size. One of the themes of his wide-ranging work was the use of specific stochastic processes, Markov chains, generalized birth-death processes and others, for the detailed interpretation of biological or physical science data; his final post was as Professor of Biomathematics. An aspect of both Fisher’s and Bartlett’s work, which I must admit I admire, is a total lack of concern with formal mathematical regularity conditions which, of course subject to due care, typically contribute little to understanding. Professor Efron in his theoretical papers manages with great skill and panache to combine careful mathematical discussion with judicious statistical emphasis. Special stochastic models probably tend nowadays to be treated much more often by computer simulation than by mathematical analysis of the differential equations involved with obvious gains and some loss, especially when the model is intended to lead to semi-qualitative understanding of an empirical phenomenon. The aethos is rather different from the use of families of regression-type models and even more so from that of machine learning, the latter achieving very specific purposes often with high effectiveness, as Professor Efron’s discussion so elegantly illustrates. But how useful are they as a guide to deeper interpretation or for broader context prediction? I reemphasize my admiration for Professor Efron’s work and for this paper.
A broad review is given of the role of statistical concepts in the design of studies, in various aspects of data collection and definition and especially in the analysis and interpretation of data.Outline examples are given from various fields of application with some emphasis on epidemiology and medical statistics.The role of probability in various stages of investigation is outlined, in particular its place in the assessment of uncertainty in the conclusions.
Comparison of two treatments in matched pairs is a powerful general method for improving precision. When the outcome is binary the formulation in terms of logistic comparisons leads to an analysis in which concordant pairs, that is pairs in which both members show the same outcome, are discarded. The present paper discusses a number of conceptual aspects of this including a comparison with a linear in probabilities formulation, the relation between logistic parameters in different designs and in particular some new efficiency comparisons. Some emphasis is based on new relations between estimated effects derived from different formulations and on comparative calculations of asymptotic efficiency.
Background: Biomarkers of inflammation predictive of cystic fibrosis (CF) disease outcomes would increase the power of clinical trials and contribute to better personalization of clinical assessments. A representative patient cohort would improve searching for believable, generalizable, reproducible and accurate biomarkers. Methods: We recruited patients from Mountain West CF Consortium (MWCFC) care centers for prospective observational study of sputum biomarkers of inflammation. After informed consent, centers enrolled randomly selected patients with CF who were clinically stable sputum producers, 12years of age and older, without previous organ transplantation. Results: From December 8, 2014 through January 16, 2016, we enrolled 114 patients (53 male) with CF with continuing data collection. Baseline characteristics included mean age 27 years (SD = 12), 80% predicted forced expiratory volume in 1s (SD = 23%), 1.0 prior year pulmonary exacerbations (SD = 1.2), home elevation 328 m (SD=112) above sea level. Compared with other patients in the US CF Foundation Patient Registry (CFFPR) in 2014, MWCFC patients had similar distribution of sex, age, lung function, weight and rates of exacerbations, diabetes, pancreatic insufficiency, CF-related arthropathy and airway infections including methicillin-sensitive or -resistant Staphylococcus aureus, Pseudomonas aeruginosa, Burkholderia cepacia complex, fungal and non-tuberculous Mycobacteria infections. They received CF-specific treatments at similar frequencies. Conclusions: Randomly-selected, sputum-producing patients within the MWCFC represent sputum-producing patients in the CFFPR. They have similar characteristics, lung function and frequencies of pulmonary exacerbations, microbial infections and use of CF-specific treatments. These findings will plausibly make future interpretations of quantitative measurements of inflammatory biomarkers generalizable to sputum-producing patients in the CFFPR.
International Statistical ReviewVolume 87, Issue 2 p. 443-443 Book Review Foundations of Info-Metrics, Amos Golan, Oxford University Press, 2018, xx + 465 pages, $99.00, hardcover, ISBN: 978-0-199-34953-1 D. R. Cox, D. R. Cox david.cox@nuffield.ox.ac.uk Nuffield College, Oxford OX1 1NF, UKSearch for more papers by this author D. R. Cox, D. R. Cox david.cox@nuffield.ox.ac.uk Nuffield College, Oxford OX1 1NF, UKSearch for more papers by this author First published: 13 August 2019 https://doi.org/10.1111/insr.12341Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinkedInRedditWechat Volume87, Issue2August 2019Pages 443-443 RelatedInformation
The analysis of binary response data commonly uses models linear in the logistic transform of probabilities. This paper considers some of the advantages and disadvantages of simple least-squares estimates based on a linear representation of the probabilities themselves, this in particular sometimes allowing a more direct empirical interpretation of underlying parameters. A sociological study is used in illustration.
Summary Methods for the analysis of large numbers of p ‐values, observed levels of significance, are reviewed with some emphasis placed on the Rényi decomposition hinging on the relation with independent exponentially distributed random variables. Some extensions are described. An empirical example is used in illustration, and simulation results examining potential complications are outlined.
Interactions in the airway ecology of cystic fibrosis may alter organism persistence and clinical outcomes. Better understanding of such interactions could guide clinical decisions. We used generalized estimating equations to fit logistic regression models to longitudinal 2-year patient cohorts in the Cystic Fibrosis Foundation Patient Registry, 2003 to 2011, in order to study associations between the airway organisms present in each calendar year and their presence in the subsequent year. Models were adjusted for clinical characteristics and multiple observations per patient. Adjusted models were tested for sensitivity to cystic fibrosis-specific treatments. The study included 28,042 patients aged 6 years and older from 257 accredited U.S. care centers and affiliates. These patients had produced sputum specimens for at least two consecutive years that were cultured for methicillin-sensitive Staphylococcus aureus, methicillin-resistant S. aureus, Pseudomonas aeruginosa, Burkholderia cepacia complex, Stenotrophomonas maltophilia, Achromobacter xylosoxidans, and Candida and Aspergillus species. We analyzed 99.8% of 538,458 sputum cultures from the patients during the study period. Methicillin-sensitive S. aureus was negatively associated with subsequent Paeruginosa. Paeruginosa was negatively associated with subsequent B. cepacia complex, Axylosoxidans, and Smaltophilia. Bcepacia complex was negatively associated with the future presence of all bacteria studied, as well as with that of Aspergillus species. Paeruginosa, B. cepacia complex, and S. maltophilia were each reciprocally and positively associated with Aspergillus species. Independently of patient characteristics, the organisms studied interact and alter the outcomes of treatment decisions, sometimes in unexpected ways. By inhibiting P. aeruginosa, methicillin-sensitive S. aureus may delay lung disease progression. Paeruginosa and B. cepacia complex may inhibit other organisms by decreasing airway biodiversity, potentially worsening lung disease.
Recently, Cox and Battey (2017 Proc. Natl Acad. Sci. USA 114 , 8592–8595 ( doi:10.1073/pnas.1703764114 )) outlined a procedure for regression analysis when there are a small number of study individuals and a large number of potential explanatory variables, but relatively few of the latter have a real effect. The present paper reports more formal statistical properties. The results are intended primarily to guide the choice of key tuning parameters.