Reserve is a physiological capacity used under demanding situations. The concept was developed to account for the discrepancy between pathology and clinical manifestation. In neuroscience, motor, brain and cognitive reserves are abstract measures, conceptually defined yet elusive to quantify. Reserve is indirectly assessed using proxies such as years of education and brain volume, limiting its utility. Moreover, the dichotomy in definitions of cognitive and motor reserves is artificial, as daily function requires an intricate network of connections between these domains. Here, we assessed the validity of a newly developed graded motor cognitive 'stress test' to quantify the combined motor and cognitive reserve (MCR). The study included 144 participants (ages between 18 and 85, 50% women) with a range of reserve capacities (i.e. healthy young and older adults and individuals with Parkinson's disease, Alzheimer's disease, dementia with Lewy bodies and mild cognitive impairment). The assessment included walking on a treadmill while negotiating motor and cognitive challenges delivered using virtual reality. To establish an MCR index score, we used a semi-supervised machine learning algorithm. The model includes performance measures from completing the stress test and measures obtained from wearable sensors used during the test. Validation of the proposed MCR index was examined through: (i) model face validity-reflecting decline of performance as challenge increased; (ii) known-groups validity-classification of scores according to neurological status; (iii) construct validity (convergent)-association with common MCRs proxies as well as MRI-derived regional brain volumes. The model's face validity revealed decreased performance with increased motor and cognitive challenges (both domains P < 0.001). The index accurately discriminated between healthy controls and those diagnosed with neurological conditions with an area under the curve of 0.89 [95% CI: 0.79-0.99] which was significantly higher than all other commonly used proxies. Statistically significant Spearman's ρ correlations were observed with all commonly used motor and cognitive proxies (0.56 ≤ r ≤ 0.79, after multiplicity correction all P < 0.05), reflecting construct validity. In addition, statistically significant correlations were observed between the MCR index and whole-brain grey matter and white matter volumes (r = 0.63 and 0.55), as well as the pre-defined left and right caudate nucleus (r = 0.56 and 0.68) and inferior-frontal gyrus (r = 0.47 and 0.58). This proof-of-concept study shows that the novel MCR index is valid, with high sensitivity to neurological deficits and is able to quantify reserve on an individual level. This new innovative tool can assist in screening for motor cognitive deficits and potentially, for predicting motor and cognitive decline associated with neurodegenerative disease.
This study addresses the challenges of inference following selection in fields like clinical trials, genome-wide association studies, and functional magnetic resonance imaging, where traditional methods like simultaneous confidence intervals (CIs) might be too conservative. We introduce an improved false coverage-statement rate controlling CIs, when the selection is done by passing a threshold in a certain direction. The CIs for the selected parameters are similar to those proposed by Benjamini and Yekutieli (2005) on the inward end, and to the standard nonadjusted CIs on the outward end. The centre of the suggested CI is a shrunk estimator of the selected parameter. This simple improvement is uniformly better for the one-directional selection, and we also suggest how to apply it for a two-directional selection. We prove that the false coverage-statement control for independent estimators and provide simulation evidence for its robustness under dependency.
Confidence intervals (CIs) are instrumental in statistical analysis, providing a range estimate of the parameters. In modern statistics, selective inference is common, where only certain parameters are highlighted. However, this selective approach can bias the inference, leading some to advocate for the use of CIs over p-values. To increase the flexibility of confidence intervals, we introduce direction-preferring CIs, enabling analysts to focus on parameters trending in a particular direction. We present these types of CIs in two settings: First, when there is no selection of parameters; and second, for situations involving parameter selection, where we offer a conditional version of the direction-preferring CIs. Both of these methods build upon the foundations of Modified Pratt CIs, which rely on non-equivariant acceptance regions to achieve longer intervals in exchange for improved sign exclusions. We show that for selected parameters out of m > 1 initial parameters of interest, CIs aimed at controlling the false coverage rate, have higher power to determine the sign compared to conditional CIs. We also show that conditional confidence intervals control the marginal false coverage rate (mFCR) under any dependency.
CONTEXT:Change in ability realization reflects the main contribution of rehabilitation to improvement in the performance of daily activities after spinal cord lesions (SCL). OBJECTIVE:To adapt a Spinal Cord Ability Realization Measurement Index (SCI-ARMI) formula to the new Spinal Cord Independence Measure version 4 (SCIM4). METHODS:Using data from 156 individuals for whom American Spinal Injury Association Motor Score (AMS) and SCIM4 scores were collected, we obtained an estimate for the highest possible SCIM4 given the patient's AMS value, using the 95th percentile of SCIM4 values at discharge from rehabilitation (SCIM95) for patients with any given AMS at discharge. We used the statistical software environment R to implement the quantile regression method for linear and quadratic formulas. We also compared the computed model with the SCIM95 model obtained using data from the present study group, positioned in the SCIM95 formula developed for SCIM3. RESULTS:The coefficients of the computed SCIM95 formula based on SCIM4 scores were statistically non-significant, which hypothetically reflects the small sample relative to the goal of estimating SCIM4 95th percentile. Predicting the ability using SCIM4 scores positioned in the SCIM95 formula used for SCIM3, however, yielded SCIM95 values, which are very close to those of the new SCIM95 formula (Mean difference 2.16, 95% CI = 1.45, 4.90). CONCLUSION:The SCI-ARMI formula, which is based on the SCIM95 formula developed for SCIM3, is appropriate for estimating SCI-ARMI at present, when SCIM4 scores are available. When sufficient additional data accumulates, it will be appropriate to introduce a modified SCI-ARMI formula.
After 38 years of operational cloud seeding for rain enhancement in northern Israel, the Israel 4 experiment was conducted to reassess its effect on rainfall and provide a basis to evaluate its utility. Operational seeding started after two randomized experiments, the second ending in 1976, found a large and statistically significant effect of cloud seeding on rainfall. Observational studies in later years raised doubts as to the magnitude of the effect, possibly because of chang-ing climatological conditions. A carefully designed randomized experiment was conducted from 2013 to 2020. A unique feature of the design was the use of forecast rainfall on target, rather than rainfall in an unaffected area, as a control variate to attenuate variability. The Israel 4 experiment was stopped a year earlier than planned, because the result was disap-pointing: a 1.8% increase, p value 5 0.4, and 95% confidence interval of (211%, 16%). These results led to a decision by the Israel Water Authority to stop operational seeding.SIGNIFICANCE STATEMENT: The recent cloud seeding experiment in northern Israel did not show a significant rainfall increase}unlike the sequence of seeding experiments conducted in Israel in the previous century.
The utility of mouse and rat studies critically depends on their replicability in other laboratories. A widely advocated approach to improving replicability is through the rigorous control of predefined animal or experimental conditions, known as standardization. However, this approach limits the generalizability of the findings to only to the standardized conditions and is a potential cause rather than solution to what has been called a replicability crisis. Alternative strategies include estimating the heterogeneity of effects across laboratories, either through designs that vary testing conditions, or by direct statistical analysis of laboratory variation. We previously evaluated our statistical approach for estimating the interlaboratory replicability of a single laboratory discovery. Those results, however, were from a well-coordinated, multi-lab phenotyping study and did not extend to the more realistic setting in which laboratories are operating independently of each other. Here, we sought to test our statistical approach as a realistic prospective experiment, in mice, using 152 results from 5 independent published studies deposited in the Mouse Phenome Database (MPD). In independent replication experiments at 3 laboratories, we found that 53 of the results were replicable, so the other 99 were considered non-replicable. Of the 99 non-replicable results, 59 were statistically significant (at 0.05) in their original single-lab analysis, putting the probability that a single-lab statistical discovery was made even though it is non-replicable, at 59.6%. We then introduced the dimensionless "Genotype-by-Laboratory" (GxL) factor-the ratio between the standard deviations of the GxL interaction and the standard deviation within groups. Using the GxL factor reduced the number of single-lab statistical discoveries and alongside reduced the probability of a non-replicable result to be discovered in the single lab to 12.1%. Such reduction naturally leads to reduced power to make replicable discoveries, but this reduction was small (from 87% to 66%), indicating the small price paid for the large improvement in replicability. Tools and data needed for the above GxL adjustment are publicly available at the MPD and will become increasingly useful as the range of assays and testing conditions in this resource increases.
Linking scalp electroencephalography (EEG) signals and spontaneous firing activity from deep nuclei in humans is not trivial. To examine this, we analyzed simultaneous recordings of scalp EEG and unit activity in deeply located sites recorded overnight from patients undergoing pre-surgical invasive monitoring. We focused on modeling the within-subject average unit activity of two medial temporal lobe areas: amygdala and hippocampus. Linear regression model correlates the units' average firing activity to spectral features extracted from the EEG during wakefulness or non-REM sleep. We show that changes in mean firing activity in both areas and states can be estimated from EEG (Pearson r > 0.2, p≪0.001). Region specificity was shown with respect to other areas. Both short- and long-term fluctuations in firing rates contributed to the model accuracy. This demonstrates that scalp EEG frequency modulations can predict changes in neuronal firing rates, opening a new horizon for non-invasive neurological and psychiatric interventions.
We provide commentary on the paper by Willi Maurer, Frank Bretz, and Xiaolei Xun entitled, "Optimal test procedures for multiple hypotheses controlling for the familywise expected loss." The authors provide an excellent discussion of the multiplicity problem in clinical trials and propose a novel approach based on a decision-theoretic framework that incorporates loss functions that can vary across multiple hypotheses in a family. We provide some considerations for the practical use of the authors' proposed methods as well as some alternative methods that may also be of interest in this setting.
We discuss three issues. In the first part, we discuss the criteria emphasized by Maurer, Bretz, and Xun, warning that it modifies the per comparison error rate that does not address the concerns raised by multiple testing. In the second part, we strengthen the optimality results developed in the paper, based on our recent results. In the third part, we highlight the potentially important role that the use of weights may have in practice and discuss the difficulties in assigning weights that convey the importance in the gain and loss functions, especially as it pertains to multiple endpoints.
Stress tests, e.g., the cardiac stress test, are standard clinical screening tools aimed to unmask clinical pathology. As such stress tests indirectly measure physiological reserves. The term reserve has been developed to account for the dis-junction, often observed, between pathology and clinical manifestation. It describes a physiological capacity that is utilized in demanding situations. However, developing a new and reliable stress test based screening tool is complex, prolonged, and relies extensively on domain knowledge. We propose a novel distributional-free machine-learning framework, the Stress Test Performance Scoring (STEPS) framework, to model expected performance in a stress test. A performance scoring function is trained with measures taken during the performance in a given task while exploiting information regarding the stress test set-up and subjects' medical state. Multiple ways of aggregating performance scores at different stress levels are suggested and are examined with an extensive simulation study. When applied to a real-world data example, an AUC of 84.35[95%CI: 70.68 - 95.13] was obtained for the STEPS framework to distinguish subjects with neurodegeneration from controls. In summary, STEPS improved screening by exploiting existing domain knowledge and state-of-the-art clinical measures. The STEPS framework can ease and speed up the production of new stress tests.
Systematic reviews and meta-analyses are important tools for synthesizing evidence from multiple studies. They serve to increase power and improve precision, in the same way that large studies can do, but also to establish the consistency of effects and replicability of results across studies. In this work we propose statistical tools to quantify replicability of effect signs (or directions) and their consistency. We suggest that these tools accompany the fixed-effect or random-effects meta-analysis, and we show that they convey important information for the assessment of the intervention under investigation. We motivate and demonstrate our approach and its implications by examples from systematic reviews from the Cochrane Library. Our tools make no assumptions on the distribution of the true effect sizes, so their inferential guarantees continue to hold even if the assumptions of the fixed-effect or random-effects models do not hold. We also develop a version of this tool under the fixed-effect assumption for cases where it is crucial and justified.
BACKGROUND:One in ten newborn children is born prematurely. The elongated length of stay (LOS) of these children in the Neonatal Intensive Care Unit (NICU) has important implications on hospital occupancy figures, healthcare and management costs, as well as the psychology of parents. In order to allow accurate planning and resource allocation, this study aims to create a generalizable and robust model to predict the NICU LOS of preterm newborns. METHODS:Data were collected from a large tertiary center NICU between 2011 and 2018 and relates to 5,362 newborns. The selected model was externally validated using a data set of 8,768 newborns from another tertiary center NICU. This report compares several models, such as Random Forest (RF), quantile RF, and other feature selection methods, including LASSO and AIC step-forward selection. In addition, a novel step-forward selection based on False Discovery Rate (FDR) for quantile regression is presented and evaluated. RESULTS:A high-orderquantile regression model for predicting preterm newborns' LOS that uses only four features available at birth had more attractive properties than other richer ones. The model achieved a Mean Absolute Error (MAE) of 6.26 days on the internal validation set (average LOS 27.04) and an MAE of 6.04 days on the external validation set (average LOS 29.32). The suggested model surpassed the accuracy obtained by models in the literature. It is shown empirically that the FDR-based selection has better properties than the AIC-based step-forward selection approach. CONCLUSION:This paper demonstrates a process to create a predictive model for NICU LOS in preterm newborns, where each step is reasoned. We obtain a simple and robust model for NICU LOS prediction, which achieves far better results than the current model used for financing NICUs. Utilizing this model, we have created an easy-to-use online web application to ease parents' worries and to assist NICU management: https://tzviel.shinyapps.io/calcuLOS.
Mathematical and statistical models have played an important role in the analysis of data from COVID-19. They are important for tracking the progress of the pandemic, for understanding its spread in the population, and perhaps most significantly for forecasting the future course of the pandemic and evaluating potential policy options. This article describes the types of models that were used by research teams in Israel, presents their assumptions and basic elements, and illustrates how they were used, and how they influenced decisions. The article grew out of a "modelists' dialog" organized by the Israel National Institute for Health Policy Research with participation from some of the leaders in the local modeling effort.
Objective: To examine the fourth version of the Spinal Cord Independence Measure for reliability and validity. Design: Partly blinded comparison with the criterion standard Spinal Cord Independence Measure III, and between examiners and examinations. Setting: A multicultural cohort from 19 spinal cord injury units in 11 countries. Participants: A total of 648 patients with spinal cord injury. Intervention: Assessment with Spinal Cord Independence Measure (SCIM IV) and Spinal Cord Independence Measure (SCIM III) on admission to inpatient rehabilitation and before discharge. Main outcome measures: SCIM IV interrater reliability, internal consistency, correlation with and difference from SCIM III, and responsiveness. Results: Total agreement between examiners was above 80% on most SCIM IV tasks. All Kappa coefficients were above 0.70 and statistically significant (P<. 001). Pearson's coefficients of the correlation between the examiners were above 0.90, and intraclass correlation coefficients were above 0.90. Cronbach's alpha was above 0.96 for the entire SCIM IV, above 0.66 for the subscales, and usually decreased when an item was eliminated. Reliability values were lower for the subscale of respiration and sphincter management, and on admission than at discharge. SCIM IV and SCIM III mean values were very close, and the coefficients of Pearson correlation between them were 0.91-0.96 (P<. 001). The responsiveness of SCIM IV was not significantly different from that of SCIM III in most of the comparisons. Conclusions: The validity, reliability, and responsiveness of SCIM IV, which was adjusted to assess specific patient conditions or situations that SCIM III does not address, and which includes more accurate definitions of certain scoring criteria, are very good and quite similar to those of SCIM III. SCIM IV can be used for clinical and research trials, including international multi-center studies, and its group scores can be compared with those of SCIM III. (C) 2021 The American Congress of Rehabilitation Medicine. Published by Elsevier Inc. All rights reserved.
Experimentation with mouse and rat models has become a central strategy for discovering mammalian gene function, and for preclinical testing of pharmacological treatments, yet the utility of any findings critically depends on their replicability in other laboratories. In previous publications we proposed a statistical approach for estimating the inter-laboratory replicability of novel discoveries made in a single laboratory. We demonstrated that previous phenotyping results from multi-lab databases can be used to derive a Genotype-by-Lab (GxL) adjustment factor to greatly enhance the replicability of the single-lab findings, for similarly measured phenotypes, even before making the effort of replicating these finding in additional laboratories. This demonstration, however, still raised several important questions that could only be answered by an additional large-scale prospective experiment: 1) Does GxL-adjustment work in single-lab experiments that were not intended to be standardized across laboratories, and with genotypes that were not included in the previous experiments? And 2) Can it be used to adjust the results of pharmacological experiments? We investigated these questions by attempting to replicate, across three laboratories, results from five single-lab studies in the Mouse Phenome Database (MPD), offering 212 comparisons, including 60 involving a pharmacological treatment: 18 mg/kg/day fluoxetine. In addition, we define and use a dimensionless GxL factor, by dividing the GxL variance by the standard deviation between animals within groups, as a more robust vehicle to transfer the adjustment from the multi-lab analysis to very different labs and genotypes. For genotype comparisons, GxL-adjustment reduced the rate of non-replicable discoveries from 60% to 12%, for the price of reducing the power to make replicable discoveries from 87% to 66%. In absolute numbers, the adjustment prevented 23 non-replicable discoveries for the price of missing only three replicated ones. Tools and data needed for deployment of this method across other mouse experiments are publicly available in MPD. Our results further point at some phenotypes as more prone to produce non-replicable results, while others, known to be more difficult to measure, are as likely to produce replicable results (once adjusted) such as the physiological measure, body weight.
We propose a method of testing a shift between mean vectors of two multivariate Gaussian random variables in a high-dimensional setting incorporating the possible dependency and allowing $$p > n$$ . This method is a combination of two well-known tests: the Hotelling test and the Simes test. The tests are integrated by sampling several dimensions at each iteration, testing each using the Hotelling test, and combining their results using the Simes test. We prove that this procedure is valid asymptotically. This procedure can be extended to handle non-equal covariance matrices by plugging in the appropriate extension of the Hotelling test. Using a simulation study, we show that the proposed test is advantageous over state-of-the-art tests in many scenarios and robust to violation of the Gaussian assumption.
The value of hypothesis testing, and the frequent misinterpretation of p-values as a cornerstone of statistical methodology, continues to be debated. In 2019, the President of the American Statistical Association (ASA) convened a Task Force to write a succinct statement about the use of statistical methods in scientific studies, specifically hypothesis tests and p-values, and their connection to replicability.
Replicability of results has been a gold standard in science and should remain so, but concerns about lack of it have increased in recent years. Transparency, good design, and reproducible computing and data analysis are prerequisites for replicability. Adopting appropriate statistical methodologies is another identified one, yet which methodologies can be used to enhance replicability of results from a single study remains controversial. Whereas the p-value and statistical significance are carrying most of the blame, this article argues that addressing selective inference is a missing statistical cornerstone of enhancing replicability. I review the manifestation of selective inference and the available ways to address it. I also discuss and demonstrate whether and how selective inference is addressed in many fields of science, including the attitude of leading scientific publications as expressed in their recent editorials. Most notably, selective inference is attended when the number of potential findings from which the selection takes place is in the thousands, but it is ignored when 'only' dozens and hundreds of potential discoveries are involved. As replicability, and its closely related concept of generalizability, can only be assessed by actual replication attempts, the question of how to make replication an integral part of the regular scientific work becomes crucial. I outline a way to ensure that some replication effort will be an inherent part of every study. This approach requires the efforts and cooperation of all parties involved: scientists, publishers, granting agencies, and academic leaders.