Machine learning classification systems are susceptible to poor performance when trained with incorrect ground truth labels, even when data is well-curated by expert annotators. As machine learning becomes more widespread, it is increasingly imperative to identify and correct mislabeling to develop more powerful models. In this work, we motivate and describe Adaptive Label Error Detection (ALED), a novel method of detecting mislabeling. ALED extracts an intermediate feature space from a deep convolutional neural network, denoises the features, models the reduced manifold of each class with a multidimensional Gaussian distribution, and performs a simple likelihood ratio test to identify mislabeled samples. We show that ALED has markedly increased sensitivity, without compromising precision, compared to established label error detection methods, on multiple medical imaging datasets. We demonstrate an example where fine-tuning a neural network on corrected data results in a 33.8
Diabetes affects one in ten adults worldwide, yet clinicians lack tools to simultaneously predict which complications a given patient is most likely to develop in the near term. Existing risk models typically address single complications, were developed in specialized cohorts, and use heterogeneous time horizons, limiting their utility for individualized clinical decision-making. Here, we show that machine learning models trained on nationwide U.S. administrative claims data for 400,400 adults with newly diagnosed type 1 or type 2 diabetes can accurately predict the monthly risk of nine acute and chronic complications (cardiovascular disease, cerebrovascular disease, peripheral vascular disease, nephropathy, neuropathy, retinopathy, and hypoglycemic and hyperglycemic crises) with predictions that update dynamically as new clinical data become available. Using Regularized Logistic Regression (GLMnet) and Extreme Gradient Boosting (XGBoost) with walk-forward temporal validation, we demonstrate strong discrimination in external validation against an independent electronic health record cohort (areas under the receiver operating characteristic curve 0.77–0.85 across complications for the best-performing model). The Diabetes Complications Risk Calculator therefore provides encounter-specific near-term risk estimates for multiple diabetes complications using routinely available clinical data, offering a potential framework for personalized risk stratification that could support shared decision-making and population health management across healthcare settings. The generalizability of the Diabetes Complications Risk Calculator to other cohorts will need to be tested. Here the researchers developed novel machine learning models that use routine clinical data to predict the risks of the nine common acute and chronic diabetes complications with predictions that update dynamically as new clinical data become available, with validation across two independent U.S. healthcare populations.
Neural organoids (NOs) have emerged as important tissue engineering models for microphysiological systems, brain sciences, and biocomputing. Establishing reliable relationships between stimulation and recording traces of electrical activity is essential for monitoring the functionality of NOs, especially in paradigms such as neural plasticity, learning, or stimulus discrimination. While researchers have demonstrated neuromodulation in NOs, they have primarily used 2D microelectrode arrays (MEAs) with limited access to the entire 3D contour of the NOs. Here, we report neuromodulation using tiny mimics of macroscale EEG caps, or shell MEAs. Specifically, we observe that stimulating current within a specific range (20 to 30 µA) induced a statistically significant increase in neuron firing rate when comparing the activity 5 s before and after stimulation. We detect neuromodulatory behavior using both three- and 16-electrode shells and generated 3D spatiotemporal maps of neuromodulatory activity around the entire surface of the NO. Our studies demonstrate a methodology for investigating 3D spatiotemporal neuromodulation in organoids of broad relevance to biomedical engineering models of neural functionality, plasticity, and learning.
A number of domains in biomedical research use data with a large number of predictors all representing the same type of measurement. Often, an important summary is the within-person distribution of these predictors. Here we focus on settings where the mean relationship between outcome and predictors is fully captured by this distribution and, more generally, on problems where the goal is to learn a mapping that is invariant under permutations of the input vector. We compare unstructured neural networks, which do not explicitly incorporate the permutation invariance property, versus networks that we call ordered predictors neural networks. We show in simulations that the unstructured deep learning approach can yield higher prediction errors, compared to the approach that explicitly leverages the invariance to simplify the learning task. Additionally, in the context of neural Bayes estimation, in which neural networks are used to construct point estimators, we show that ordered predictors neural networks can yield substantially more precise estimators. We therefore recommend that, when permutation invariance is known or suspected to hold, investigators use a learning or statistical modeling approach that can leverage the invariance, rather than an unstructured deep learning approach.
Here, to enable researchers to more fully harness the collective discovery potential of multiomic data in the public domain, we have assembled gene-level transcriptomic data from ~200 studies of neocortical development and in vitro models. Applying joint matrix decomposition to mouse, macaque and human data, we define transcriptome dynamics that are conserved across neocortical neurogenesis and identify a program that emerges in ventricular progenitors, is later expressed in neurogenic outer, or basal, radial glia of primates, but is limited to gliogenic precursors in the rodent. Decomposition of adult human neocortical data identified layer-specific signatures in excitatory neurons, enabling the charting of their developmental emergence and protracted maturation, which is in stark contrast to the early peaking expression of layer-defining transcription factors. Interrogation of data from cerebral organoids demonstrated that, although broad elements of in vivo development are recapitulated in vitro, many layer-specific transcriptomic programs in neuronal maturation are absent. We invite cell biologists without coding expertise to use NeMO Analytics in their research and to fuel it with their own emerging data at nemoanalytics.org/landing/neocortex .
BackgroundUnderstanding individual variability in response to interventions is essential for developing personalized treatment strategies. In rare and clinically heterogeneous conditions like primary progressive aphasia (PPA), predicting treatment response is particularly challenging due to varying clinical manifestations. In this study, we aimed to identify and analyze predictors of individual language response to transcranial direct current stimulation (tDCS) of the left inferior frontal gyrus (IFG), using a novel, robust analytic approach focused on treatment effect heterogeneity.MethodsWe compared the ability of predicting individual effect (active vs sham tDCS during 20-minute sessions on weekdays for 3 weeks; active: 2 mA current across electrodes; sham: current ramped down after 30 seconds), using demographic and clinical patient characteristics (eg, PPA variant and disease progression, baseline language performance) or volumetric fMRI data versus functional connectivity (from resting-state fMRI) in the cohort of 36 patients.ResultsFunctional connectivity alone had the highest predictive value for outcomes, explaining 62% of the variance of the tDCS effect in generalization (semantic fluency) and 75% of the main outcome (written naming), contrasted with <15% (for semantic fluency) and <23% (for written naming) of variance predicted by demographic and clinical patient characteristics or volumetric data. Patients with higher baseline functional connectivity within the left IFG (between pars opercularis and pars triangularis) were most likely to benefit from tDCS both in generalization (semantic fluency) as well as in the main outcome (written naming). In addition, patients with higher baseline FC between the middle temporal pole and superior temporal gyrus, were most likely to show generalization effects of tDCS.ConclusionsThe present study showcases the importance of a baseline functional connectivity scan in predicting tDCS outcomes, and points toward a precision medicine approach in neuromodulation studies. The study has important implications for clinical trials and practice, providing a statistical method that addresses heterogeneity in patient populations and allowing accurate prediction and enrollment of those who will most likely benefit from specific interventions.
Ordinal data is widely prevalent in clinical and other domains, yet there is a lack of both modern, machine-learning based methods and publicly available software to address it. In this paper, we present a model-agnostic method of ordinal classification, which can apply any non-ordinal classification method in an ordinal fashion. We also provide an open-source implementation of these algorithms, in the form of a Python package. We apply these models on multiple real-world datasets to show their performance across domains. We show that they often outperform non-ordinal classification methods, especially when the number of datapoints is relatively small or when there are many classes of outcomes. This work, including the developed software, facilitates the use of modern, more powerful machine learning algorithms to handle ordinal data.
Accurate assessment of Alzheimer's disease (AD) using neuroimaging is important for understanding disease progression and supporting clinical research. Structural MRI (sMRI) is widely used for this purpose, but most deep learning approaches primarily rely on intensity-based features, which may not fully capture subtle morphological variations. Jacobian determinant maps (JSM) provide complementary information by describing localized brain deformations; however, conventional multimodal fusion strategies such as early or late fusion may not fully capture the interactions between these modalities. To address this, we explore a cross-attention fusion approach that explicitly models the relationship between sMRI intensity and JSM-derived deformation features for AD classification. Using data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), we compare cross-attention with pairwise self-attention and four baseline fusion methods. The proposed approach achieved mean ROC-AUC scores of 0.903 (0.033) for AD vs. cognitively normal (CN) and 0.692 (0.061) for mild cognitive impairment (MCI) vs. CN under five-fold cross-validation. The proposed approach achieved the best performance among fusion strategies and achieved higher or comparable performance with only 1.56 million parameters. These findings suggest that cross-attention fusion may serve as a potentially useful approach for integrating structural and deformation-based MRI features in AD classification tasks.
Recent advances in spatially resolved single-omic and multi-omics technologies have led to the emergence of computational tools to detect and predict spatial domains. Additionally, histological images and immunofluorescence (IF) staining of proteins and cell types provide multiple perspectives and a more complete understanding of tissue architecture. Here, we introduce Proust, a scalable tool to predict discrete domains using spatial multi-omics data by combining the low-dimensional representation of biological profiles based on graph-based contrastive self-supervised learning. Our scalable method integrates multiple data modalities, such as RNA, protein, and H&E images, and predicts spatial domains within tissue samples. Through the integration of multiple modalities, Proust consistently demonstrates enhanced accuracy in detecting spatial domains, as evidenced across various benchmark data sets and technological platforms.
Early diagnosis of Alzheimer's disease (AD) is critical for intervention before irreversible neurodegeneration occurs. Structural MRI (sMRI) is widely used for AD diagnosis, but conventional deep learning approaches primarily rely on intensity-based features, which require large datasets to capture subtle structural changes. Jacobian determinant maps (JSM) provide complementary information by encoding localized brain deformations, yet existing multimodal fusion strategies fail to fully integrate these features with sMRI. We propose a cross-attention fusion framework to model the intrinsic relationship between sMRI intensity and JSM-derived deformations for AD classification. Using the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, we compare cross-attention, pairwise self-attention, and bottleneck attention with four pre-trained 3D image encoders. Cross-attention fusion achieves superior performance, with mean ROC-AUC scores of 0.903 (+/-0.033) for AD vs. cognitively normal (CN) and 0.692 (+/-0.061) for mild cognitive impairment (MCI) vs. CN. Despite its strong performance, our model remains highly efficient, with only 1.56 million parameters–over 40 times fewer than ResNet-34 (63M) and Swin UNETR (61.98M). These findings demonstrate the potential of cross-attention fusion for improving AD diagnosis while maintaining computational efficiency.
Recent advances in deep learning have made it possible to predict phenotypic measures directly from functional magnetic resonance imaging (fMRI) brain volumes, sparking significant interest in the neuroimaging community. However, existing approaches, primarily based on convolutional neural networks or transformer architectures, often struggle to model the complex relationships inherent in fMRI data, limited by their inability to capture long-range spatial and temporal dependencies. To overcome these shortcomings, we introduce BrainMT, a novel hybrid framework designed to efficiently learn and integrate long-range spatiotemporal attributes in fMRI data. Our framework operates in two stages: (1) a bidirectional Mamba block with a temporal-first scanning mechanism to capture global temporal interactions in a computationally efficient manner; and (2) a transformer block leveraging self-attention to model global spatial relationships across the deep features processed by the Mamba block. Extensive experiments on two large-scale public datasets, UKBioBank and the Human Connectome Project, demonstrate that BrainMT achieves state-of-the-art performance on both classification (sex prediction) and regression (cognitive intelligence prediction) tasks, outperforming existing methods by a significant margin. Our code and implementation details are available at link.
Batch effects, undesirable sources of variability across multiple experiments, present significant challenges for scientific and clinical discoveries. Batch effects can (i) produce spurious signals and/or (ii) obscure genuine signals, contributing to the ongoing reproducibility crisis. Because batch effects are typically modeled as classical statistical effects, they often cannot differentiate between sources of variability due to confounding biases, which may lead them to erroneously conclude batch effects are present (or not). We formalize batch effects as causal effects, and introduce algorithms leveraging causal machinery, to address these concerns. Simulations illustrate that when non-causal methods provide the wrong answer, our methods either produce more accurate answers or "no answer," meaning they assert the data are inadequate to confidently conclude on the presence of a batch effect. Applying our causal methods to 27 neuroimaging datasets yields qualitatively similar results: in situations where it is unclear whether batch effects are present, non-causal methods confidently identify (or fail to identify) batch effects, whereas our causal methods assert that it is unclear whether there are batch effects or not. In instances where batch effects should be discernable, our techniques produce different results from prior art, each of which produce results more qualitatively similar to not applying any batch effect correction to the data at all. This work, therefore, provides a causal framework for understanding the potential capabilities and limitations of analysis of multi-site data.
The emergence of functional magnetic resonance imaging (fMRI) marked a significant technological breakthrough in the real-time measurement of the functioning human brain in vivo. In part because of their 4D nature (three spatial dimensions and time), fMRI data have inspired a great deal of statistical development in the past couple of decades to address their unique spatiotemporal properties. This article provides an overview of the current landscape in functional brain measurement, with a particular focus on fMRI, highlighting key developments in the past decade. Furthermore, it looks ahead to the future, discussing unresolved research questions in the community and outlining potential research topics for the future.
Functional connectivity, reflecting synchronized brain activity across distinct regions, is crucial for understanding cognitive processes. Despite the recent interest in exploring the relationship between functional connectivity and structural brain features, understanding the precise link remains challenging. We propose a novel analysis method that integrates structural factors-such as anatomical morphology summaries, voxel intensity, diffusion-weighted information, and geographic distance to explain variation in functional connectivity. Our method employs generalized additive model (GAM), leveraging region-pair or vertex-pair information, while accommodating individual subject differences in both template and subject spaces. Furthermore, we assess repeatability via the so called discriminability of subjects under our approach, quantifying the probability of similarities between measurements for the same subject versus different subjects. Utilizing data from the Human Connectome Project, we analyze brain connectivity in twin pairs and non-twin pairs to evaluate the repeatability of model-based connectivity patterns estimated via GAMs. Our findings suggest that direct structure/function regression models enhances our understanding of functional connectivity variation, providing insights into underlying mechanisms and discriminability of brain connections.
Background: There is a substantial history studying the relationship between general intelligence and the core, diagnostic symptoms of autism. Given that hierarchical models of intelligence are used in clinical practice, one one thing that remains is unclear is at which level of these hierarchical models we find associations with the magnitude of autism diagnostic symptoms, and whether associations differ between versions of clinical intelligence quotient (IQ) tests. Method: We examined associations between autism diagnostic symptom magnitude, as measured by the Autism Diagnostic Observational Schedule (ADOS), and hierarchical models of general intelligence. Because previous work using the same cohorts demonstrated that the manualized index structure of the Wechsler Intelligence Scale for Children, fifth edition (WISC-V) and Wechsler Intelligence Scale for Children, fourth edition (WISC-IV) do not correlate well with data-driven factorization of subtest results within autism samples, our measures of general intelligence included Spearman’s g and data-driven intermediate factors derived from WISC-V (N=83) and the WISC-IV (N=131) subtest performance.Results: In the WISC-V cohort, ADOS scores did not show a significant correlation with g and correlated with only one empirically-derived factor score in this sample. On the other hand, in the WISC-IV, ADOS scores were correlated with g and three out of four factor scores.Conclusions: Autism core symptom presentation is more independent of general intelligence as measured by the WISC-V than as measured by WISC-IV at both the overall (full-scale IQ) and factor levels of the hierarchy.
The advent of modern data collection and processing techniques has seen the size, scale and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the characteristics of the data which are repeatable-the aspects of the data that are able to be identified under duplicated analyses. Conflictingly, the utility of traditional repeatability measures, such as the intra-class correlation coefficient, under these settings is limited. In recent work, novel data repeatability measures have been introduced in the context where a set of subjects are measured twice or more, including: fingerprinting, rank sums and generalisations of the intra-class correlation coefficient. However, the relationships between, and the best practices among, these measures remains largely unknown. In this manuscript, we formalise a novel repeatability measure, discriminability. We show that it is deterministically linked with the intra-class correlation coefficients under univariate random effect models and has the desired property of optimal accuracy for inferential tasks using multivariate measurements. Additionally, we overview and systematically compare existing repeatability statistics with discriminability, using both theoretical results and simulations. We show that the rank sum statistic is deterministically linked to a consistent estimator of discriminability. The statistical power of permutation tests derived from these measures are compared numerically under Gaussian and non-Gaussian settings, with and without simulated batch effects. Motivated by both theoretical and empirical results, we provide methodological recommendations for each benchmark setting to serve as a resource for future analyses. We believe these recommendations will play an important role towards improving repeatability in fields such as functional magnetic resonance imaging, genomics, pharmacology and more.