BackgroundAccurate assessment of mental disorders and learning disabilities is essential for timely intervention. Machine learning and feature selection techniques have demonstrated potential in improving the accuracy and efficiency of mental health assessments. However, limited research has explored the use of large transdiagnostic datasets containing a vast number of items (exceeding 1000), as well as the application of these techniques in developing quick, question-based learning disability assessments. The goals of this study are to apply machine learning and feature selection techniques to a large transdiagnostic dataset featuring a high number of input items, and to create a tool for the streamlined creation of efficient and effective assessment using existing datasets.MethodsThis study leverages the Healthy Brain Network (HBN) dataset to develop a tool for creation of efficient and effective machine learning-based assessment of mental disorders and learning disabilities. Feature selection algorithms were applied to identify parsimonious item subsets. Modular architecture ensures straightforward application to other datasets. ResultsMachine learning models trained on the HBN data exhibited improved performance over existing assessments. Using only non-proprietary assessments did not significantly impact model performance. DiscussionThis study demonstrates the feasibility of using existing large-scale datasets for creating accurate and efficient assessments for mental disorders and learning disabilities. The performance values of the machine learning models provide estimates of the performance of the new assessments in a population similar to HBN. The trained models can be used in a new population after validation and acquiring consent of the authors of the original assessments. The modular architecture of the developed tool ensures seamless application to diverse clinical and research contexts.
AI is now a cornerstone of modern dataset analysis. In many real world applications, practitioners are concerned with controlling specific kinds of errors, rather than minimizing the overall number of errors. For example, biomedical screening assays may primarily be concerned with mitigating the number of false positives rather than false negatives. Quantifying uncertainty in AI-based predictions, and in particular those controlling specific kinds of errors, remains theoretically and practically challenging. We develop a strategy called multidimensional informed generalized hypothesis testing (MIGHT) which we prove accurately quantifies uncertainty and confidence given sufficient data, and concomitantly controls for particular error types. Our key insight was that it is possible to integrate canonical cross-validation and parametric calibration procedures within a nonparametric ensemble method. Simulations demonstrate that while typical AI based-approaches cannot be trusted to obtain the truth, MIGHT can be. We apply MIGHT to answer an open question in liquid biopsies using circulating cell-free DNA (ccfDNA) in individuals with or without cancer: Which biomarkers, or combinations thereof, can we trust? Performance estimates produced by MIGHT on ccfDNA data have coefficients of variation that are often orders of magnitude lower than other state of the art algorithms such as support vector machines, random forests, and Transformers, while often also achieving higher sensitivity. We find that combinations of variable sets often decrease rather than increase sensitivity over the optimal single variable set because some variable sets add more noise than signal. This work demonstrates the importance of quantifying uncertainty and confidence-with theoretical guarantees-for the interpretation of real-world data.
Multiple case-controlled studies have shown that analyzing fragmentation patterns in plasma cell-free DNA (cfDNA) can distinguish individuals with cancer from healthy controls. However, there have been few studies that investigate various types of cfDNA fragmentomics patterns in individuals with other diseases. We therefore developed a comprehensive statistic, called fragmentation signatures, that integrates the distributions of fragment positioning, fragment length, and fragment end-motifs in cfDNA. We found that individuals with venous thromboembolism, systemic lupus erythematosus, dermatomyositis, or scleroderma have cfDNA fragmentation signatures that closely resemble those found in individuals with advanced cancers. Furthermore, these signatures were highly correlated with increases in inflammatory markers in the blood. We demonstrate that these similarities in fragmentation signatures lead to high rates of false positives in individuals with autoimmune or vascular disease when evaluated using conventional binary classification approaches for multicancer earlier detection (MCED). To address this issue, we introduced a multiclass approach for MCED that integrates fragmentation signatures with protein biomarkers and achieves improved specificity in individuals with autoimmune or vascular disease while maintaining high sensitivity. Though these data put substantial limitations on the specificity of fragmentomics-based tests for cancer diagnostics, they also offer ways to improve the interpretability of such tests. Moreover, we expect these results will lead to a better understanding of the process-most likely inflammatory-from which abnormal fragmentation signatures are derived.
Batch effects, undesirable sources of variability across multiple experiments, present significant challenges for scientific and clinical discoveries. Batch effects can (i) produce spurious signals and/or (ii) obscure genuine signals, contributing to the ongoing reproducibility crisis. Because batch effects are typically modeled as classical statistical effects, they often cannot differentiate between sources of variability due to confounding biases, which may lead them to erroneously conclude batch effects are present (or not). We formalize batch effects as causal effects, and introduce algorithms leveraging causal machinery, to address these concerns. Simulations illustrate that when non-causal methods provide the wrong answer, our methods either produce more accurate answers or "no answer," meaning they assert the data are inadequate to confidently conclude on the presence of a batch effect. Applying our causal methods to 27 neuroimaging datasets yields qualitatively similar results: in situations where it is unclear whether batch effects are present, non-causal methods confidently identify (or fail to identify) batch effects, whereas our causal methods assert that it is unclear whether there are batch effects or not. In instances where batch effects should be discernable, our techniques produce different results from prior art, each of which produce results more qualitatively similar to not applying any batch effect correction to the data at all. This work, therefore, provides a causal framework for understanding the potential capabilities and limitations of analysis of multi-site data.
Decision forests are widely used for classification and regression tasks. A lesser known property of tree-based methods is that one can construct a proximity matrix from the tree(s), and these proximity matrices are induced kernels. While there has been extensive research on the applications and properties of kernels, there is relatively little research on kernels induced by decision forests. We construct Kernel Mean Embedding Random Forests (KMERF), which induce kernels from random trees and/or forests using leaf-node proximity. We introduce the notion of an asymptotically characteristic kernel, and prove that KMERF kernels are asymptotically characteristic for both discrete and continuous data. Because KMERF is data-adaptive, we suspected it would outperform kernels selected a priori on finite sample data. We illustrate that KMERF nearly dominates current state-of-the-art kernel-based tests across a diverse range of high-dimensional two-sample and independence testing settings. Furthermore, our forest-based approach is interpretable, and provides feature importance metrics that readily distinguish important dimensions, unlike other high-dimensional non-parametric testing procedures. Hence, this work demonstrates the decision forest-based kernel can be more powerful and more interpretable than existing methods, flying in the face of conventional wisdom of the trade-off between the two.
ABSTRACTContextual fear learning is heavily dependent on the hippocampus. Despite evidence that catecholamines contribute to contextual encoding and memory retrieval, the precise temporal dynamics of their release in the hippocampus during behavior is unknown. In addition, new animal models are required to probe the effects of altered catecholamine synthesis on release dynamics and contextual learning. Utilizing GRABNEand GRABDAsensors,in vivofiber photometry, and two new mouse models of altered locus coeruleus norepinephrine (LC-NE) synthesis, we investigate norepinephrine (NE) and dopamine (DA) release dynamics in dorsal hippocampal CA1 during contextual fear conditioning. We report that aversive foot-shock increases both NE and DA release in dorsal CA1, while freezing behavior associated with recall of fear memory is accompanied by decreased release. Partial loss of LC-NE synthesis reveals that NE release dynamics are modulated by sex. Moreover, we find that recall of recent fear memory is sensitive to both partial and complete loss of LC-NE synthesis throughout prenatal and postnatal development, similar to prior observations of mice with global loss of NE synthesis beginning postnatally. In contrast, remote recall is compromised only by complete loss of LC-NE synthesis beginning prenatally. Overall, these findings provide novel insights into the role of NE in contextual fear and the precise temporal dynamics of both NE and DA during freezing behavior, and highlight a complex relationship between genotype, sex, and NE signaling.
Significance:Fiber photometry is a widely used technique in modern behavioral neuroscience, employing genetically encoded fluorescent sensors to monitor neural activity and neurotransmitter release in awake-behaving animals, However, analyzing photometry data can be both laborious and time-consuming. Aim:We propose the FiPhA (Fiber Photometry Analysis) app, which is a general-purpose fiber photometry analysis application. The goal is to develop a pipeline suitable for a wide range of photometry approaches, including spectrally resolved, camera-based, and lock-in demodulation. Approach:FiPhA was developed using the R Shiny framework and offers interactive visualization, quality control, and batch processing functionalities in a user-friendly interface. Results:This application simplifies and streamlines the analysis process, thereby reducing labor and time requirements. It offers interactive visualizations, event-triggered average processing, powerful tools for filtering behavioral events and quality control features. Conclusions:FiPhA is a valuable tool for behavioral neuroscientists working with discrete, event-based fiber photometry data. It addresses the challenges associated with analyzing and investigating such data, offering a robust and user-friendly solution without the complexity of having to hand-design custom analysis pipelines. This application thus helps standardize an approach to fiber photometry analysis.
The K-sample testing problem involves determining whether K groups of data points are each drawn from the same distribution. Analysis of variance is arguably the most classical method to test mean differences, along with several recent methods to test distributional differences. In this paper, we demonstrate the existence of a transformation that allows K-sample testing to be carried out using any dependence measure. Consequently, universally consistent K-sample testing can be achieved using a universally consistent dependence measure, such as distance correlation and the Hilbert-Schmidt independence criterion. This enables a wide range of dependence measures to be easily applied to K-sample testing.
Decision forests are popular tools for classification and regression. These forests naturally generate proximity matrices that measure the frequency of observations appearing in the same leaf node. While other kernels are known to have strong theoretical properties such as being characteristic, there is no similar result available for decision forest-based kernels. In addition, existing approaches to independence and k-sample testing may require unfeasibly large sample sizes and are not interpretable. In this manuscript, we prove that the decision forest induced proximity is a characteristic kernel, enabling consistent independence and k-sample testing via decision forests. We leverage this to introduce kernel mean embedding random forest (KMERF), which is a valid and consistent method for independence and k-sample testing. Our extensive simulations demonstrate that KMERF outperforms other tests across a variety of independence and two-sample testing scenarios. Additionally, the test is interpretable, and its key features are readily discernible. This work therefore demonstrates the existence of a test that is both more powerful and more interpretable than existing methods, flying in the face of conventional wisdom of the trade-off between the two.
Causal inference studies whether the presence of a variable influences an observed outcome. As measured by quantities such as the "average treatment effect," this paradigm is employed across numerous biological fields, from vaccine and drug development to policy interventions. Unfortunately, the majority of these methods are often limited to univariate outcomes. Our work generalizes causal estimands to outcomes with any number of dimensions or any measurable space, and formulates traditional causal estimands for nominal variables as causal discrepancy tests. We propose a simple technique for adjusting universally consistent conditional independence tests and prove that these tests are universally consistent causal discrepancy tests. Numerical experiments illustrate that our method, Causal CDcorr, leads to improvements in both finite sample validity and power when compared to existing strategies. Our methods are all open source and available at github.com/ebridge2/cdcorr.
Deep networks and decision forests (such as random forests and gradient boosted trees) are the leading machine learning methods for structured and tabular data, respectively. Many papers have empirically compared large numbers of classifiers on one or two different domains (e.g., on 100 different tabular data settings). However, a careful conceptual and empirical comparison of these two strategies using the most contemporary best practices has yet to be performed. Conceptually, we illustrate that both can be profitably viewed as"partition and vote"schemes. Specifically, the representation space that they both learn is a partitioning of feature space into a union of convex polytopes. For inference, each decides on the basis of votes from the activated nodes. This formulation allows for a unified basic understanding of the relationship between these methods. Empirically, we compare these two strategies on hundreds of tabular data settings, as well as several vision and auditory settings. Our focus is on datasets with at most 10,000 samples, which represent a large fraction of scientific and biomedical datasets. In general, we found forests to excel at tabular and structured data (vision and audition) with small sample sizes, whereas deep nets performed better on structured data with larger sample sizes. This suggests that further gains in both scenarios may be realized via further combining aspects of forests and networks. We will continue revising this technical report in the coming months with updated results.
Deep networks and decision forests (such as random forests and gradient boosted trees) are the leading machine learning methods for structured and tabular data, respectively. Many papers have empirically compared large numbers of classifiers on one or two different domains (e.g., on 100 different tabular data settings). However, a careful conceptual and empirical comparison of these two strategies using the most contemporary best practices has yet to be performed. Conceptually, we illustrate that both can be profitably viewed as partition and vote schemes. Specifically, the representation space that they both learn is a partitioning of feature space into a union of convex polytopes. For inference, each decides on the basis of votes from the activated nodes. This formulation allows for a unified basic understanding of the relationship between these methods. Empirically, we compare these two strategies on hundreds of tabular data settings, as well as several vision and auditory settings. Our focus is on datasets with at most 10,000 samples, which represent a large fraction of scientific and biomedical datasets. In general, we found forests to excel at tabular and structured data (vision and audition) with small sample sizes, whereas deep nets performed better on structured data with larger sample sizes. This suggests that further gains in both scenarios may be realized via further combining aspects of forests and networks. We will continue revising this technical report in the coming months with updated results.
Decision forests, including random forests and gradient boosting trees, remain the leading machine learning methods for many real-world data problems, especially on tabular data. However, most of the current implementations only operate in batch mode, and therefore cannot incrementally update when more data arrive. Several previous works developed streaming trees and ensembles to overcome this limitation. Nonetheless, we found that those state-of-the-art algorithms suffer from a number of drawbacks, including low accuracy on some problems and high memory usage on others. We therefore developed the simplest possible extension of decision trees: given new data, simply update existing trees by continuing to grow them, and replace some old trees with new ones to control the total number of trees. In a benchmark suite containing 72 classification problems (the OpenML-CC18 data suite), we illustrate that our approach, Stream Decision Forest (SDF), does not suffer from either of the aforementioned limitations. On those datasets, we also demonstrate that our approach often performs as well, and sometimes even better, than conventional batch decision forest algorithm. Thus, SDFs establish a simple standard for streaming trees and forests that could readily be applied to many real-world problems.
. Random forests (RF) and deep networks (DN) are two of the most popular machine learning methods in the current scientific literature and yield differing levels of performance on different data modalities. We wish to further explore and establish the conditions and domains in which each approach excels, particularly in the context of sample size and feature dimension. To address these issues, we tested the performance of these approaches across tabular, image
Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any type of dependency structure given sufficient sample size. One major bottleneck is the testing process: because the null distribution of distance correlation depends on the underlying random variables and metric choice, it typically requires a permutation test to estimate the null and compute the p-value, which is very costly for large amount of data. To overcome the difficulty, in this paper we propose a chi-square test for distance correlation. Method-wise, the chi-square test is non-parametric, extremely fast, and applicable to bias-corrected distance correlation using any strong negative type metric or characteristic kernel. The test exhibits a similar testing power as the standard permutation test, and can be utilized for K-sample and partial testing. Theory-wise, we show that the underlying chi-square distribution well approximates and dominates the limiting null distribution in upper tail, prove the chi-square test can be valid and universally consistent for testing independence, and establish a testing power inequality with respect to the permutation test.
The $k$-sample testing problem tests whether or not $k$ groups of data points are sampled from the same distribution. Multivariate analysis of variance (MANOVA) is currently the gold standard for $k$-sample testing but makes strong, often inappropriate, parametric assumptions. Moreover, independence testing and $k$-sample testing are tightly related, and there are many nonparametric multivariate independence tests with strong theoretical and empirical properties, including distance correlation (Dcorr) and Hilbert-Schmidt-Independence-Criterion (Hsic). We prove that universally consistent independence tests achieve universally consistent $k$-sample testing and that $k$-sample statistics like Energy and Maximum Mean Discrepancy (MMD) are exactly equivalent to Dcorr. Empirically evaluating these tests for $k$-sample scenarios demonstrates that these nonparametric independence tests typically outperform MANOVA, even for Gaussian distributed settings. Finally, we extend these non-parametric $k$-sample testing procedures to perform multiway and multilevel tests. Thus, we illustrate the existence of many theoretically motivated and empirically performant $k$-sample tests. A Python package with all independence and k-sample tests called hyppo is available from this https URL.
We introduce hyppo, a unified library for performing multivariate hypothesis testing, including independence, two-sample, and k-sample testing. While many multivariate independence tests have R packages available, the interfaces are inconsistent and most are not available in Python. hyppo includes many state of the art multivariate testing procedures. The package is easy-to-use and is flexible enough to enable future extensions. The documentation and all releases are available at https://hyppo.neurodata.io.
Hydrogen peroxide (H2O2) is an endogenous molecule that plays several important roles in brain function: it is generated in cellular respiration, serves as a modulator of dopaminergic signaling, and its presence can indicate the upstream production of more aggressive reactive oxygen species (ROS). H2O2 has been implicated in several neurodegenerative diseases, including Parkinson's disease (PD), creating a critical need to identify mechanisms by which H2O2 modulates cellular processes in general and how it affects the dopaminergic nigrostriatal pathway, in particular. Furthermore, there is broad interest in selective electrochemical quantification of H2O2, because it is often enzymatically generated at biosensors as a reporter for the presence of nonelectroactive target molecules. H2O2 fluctuations can be monitored in real time using fast-scan cyclic voltammetry (FSCV) coupled with carbon-fiber microelectrodes. However, selective identification is a critical issue when working in the presence of other molecules that generate similar voltammograms, such as adenosine and histamine. We have addressed this problem by fabricating a robust, H2O2-selective electrode. 1,3-Phenylenediamine (mPD) was electrodeposited on a carbon-fiber microelectrode to create a size-exclusion membrane, rendering the electrode sensitive to H2O2 fluctuations and pH shifts but not to other commonly studied neurochemicals. The electrodes are described and characterized herein. The data demonstrate that this technology can be used to ensure the selective detection of H2O2, enabling confident characterization of the role this molecule plays in normal physiological function as well as in the progression of PD and other neuropathies involving oxidative stress.