In the dynamic hospital setting, decision support can be a valuable tool for improving patient outcomes. Data-driven inference of future outcomes is challenging in this dynamic setting, where long sequences such as laboratory tests and medications are updated frequently. This is due in part to heterogeneity of data types and mixed-sequence types contained in variable length sequences. In this work we design a probabilistic unsupervised model for multiple arbitrary-length sequences contained in hospitalization Electronic Health Record (EHR) data. The model uses a latent variable structure and captures complex relationships between medications, diagnoses, laboratory tests, neurological assessments, and medications. It can be trained on original data, without requiring any lossy transformations or time binning. Inference algorithms are derived that use partial data to infer properties of the complete sequences, including their length and presence of specific values. We train this model on data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The results are evaluated against held-out data for predicting the length of sequences and presence of Intensive Care Unit (ICU) in hospitalization bed sequences. Our method outperforms a baseline approach, showing that in these experiments the trained model captures information in the sequences that is informative of their future values.
Human whole-brain functional connectivity networks have been shown to exhibit both local/quasilocal (e.g., a set of functional sub-circuits induced by node or edge attributes) and non-local (e.g., higher-order functional coordination patterns) properties. Nonetheless, the non-local properties of topological strata induced by local/quasilocal functional sub-circuits have yet to be addressed. To that end, we proposed a homological formalism that enables the quantification of higher-order characteristics of human brain functional sub-circuits. Our results indicate that each homological order uniquely unravels diverse, complementary properties of human brain functional sub-circuits. Noticeably, the H1 homological distance between rest and motor task was observed at both the whole-brain and sub-circuit consolidated levels, which suggested the self-similarity property of human brain functional connectivity unraveled by a homological kernel. Furthermore, at the whole-brain level, the rest–task differentiation was found to be most prominent between rest and different tasks at different homological orders: (i) Emotion task (H0), (ii) Motor task (H1), and (iii) Working memory task (H2). At the functional sub-circuit level, the rest–task functional dichotomy of the default mode network is found to be mostly prominent at the first and second homological scaffolds. Also at such scale, we found that the limbic network plays a significant role in homological reconfiguration across both the task and subject domains, which paves the way for subsequent investigations on the complex neuro-physiological role of such network. From a wider perspective, our formalism can be applied, beyond brain connectomics, to study the non-localized coordination patterns of localized structures stretching across complex network fibers.
In systems and network neuroscience, many common practices in brain connectomic analysis are often not properly scrutinized. One such practice is mapping a predetermined set of sub-circuits, like functional networks (FNs), onto subjects’ functional connectomes (FCs) without adequately assessing the information-theoretic appropriateness of the partition. Another practice that goes unchallenged is thresholding weighted FCs to remove spurious connections without justifying the chosen threshold. This paper leverages recent theoretical advances in Stochastic Block Models (SBMs) to formally define and quantify the information-theoretic fitness (e.g., prominence) of a predetermined set of FNs when mapped to individual FCs under different fMRI task conditions. Our framework allows for evaluating any combination of FC granularity, FN partition, and thresholding strategy, thereby optimizing these choices to preserve the important topological features of the human brain connectomes. By applying to the Human Connectome Project with Schaefer parcellations at multiple levels of granularity, the framework showed that the common thresholding value of 0.25 was indeed information-theoretically valid for group-average FCs, despite its previous lack of justification. Our results pave the way for the proper use of FNs and thresholding methods, and provide insights for future research in individualized parcellations.
Higher-order properties of functional magnetic resonance imaging (fMRI) induced connectivity have been shown to unravel many exclusive topological and dynamical insights beyond pairwise interactions. Nonetheless, whether these fMRI-induced higher-order properties play a role in disentangling other neuroimaging modalities' insights remains largely unexplored and poorly understood. In this work, by analyzing fMRI data from the Human Connectome Project Young Adult dataset using persistent homology, we discovered that the volume-optimal persistence homological scaffolds of fMRI-based functional connectomes exhibited conservative topological reconfigurations from the resting state to attentional task-positive state. Specifically, while reflecting the extent to which each cortical region contributed to functional cycles following different cognitive demands, these reconfigurations were constrained such that the spatial distribution of cavities in the connectome is relatively conserved. Most importantly, such level of contributions covaried with powers of aperiodic activities mostly within the theta-alpha (4-12 Hz) band measured by magnetoencephalography (MEG). This comprehensive result suggests that fMRI-induced hemodynamics and MEG theta-alpha aperiodic activities are governed by the same functional constraints specific to each cortical morpho-structure. Methodologically, our work paves the way toward an innovative computing paradigm in multimodal neuroimaging topological learning. The code for our analyses is provided in https://github.com/ngcaonghi/scaffold_noise.
Functional connectomes (FCs) containing pairwise estimations of functional couplings between pairs of brain regions are commonly represented by correlation matrices. As symmetric positive definite matrices, FCs can be transformed via tangent space projections, resulting into tangent-FCs. Tangent-FCs have led to more accurate models predicting brain conditions or aging. Motivated by the fact that tangent-FCs seem to be better biomarkers than FCs, we hypothesized that tangent-FCs have also a higher fingerprint. We explored the effects of six factors: fMRI condition, scan length, parcellation granularity, reference matrix, main-diagonal regularization, and distance metric. Our results showed that identification rates are systematically higher when using tangent-FCs across the "fingerprint gradient" (here including test-retest, monozygotic and dizygotic twins). Highest identification rates were achieved when minimally (0.01) regularizing FCs while performing tangent space projection using Riemann reference matrix and using correlation distance to compare the resulting tangent-FCs. Such configuration was validated in a second dataset (resting-state).
Pulse shapes differ between neutron and gamma particles when measured with detector devices employing pulse shape discriminating (PSD) scintillators. Digitized waveforms can be used in detection systems to perform pulse shape discrimination for this application. Prior Gaussian Mixture Model (GMM) methods require access to the pulse full-waveform. Reducing the waveform to a smaller set of combined samples reduces computational cost while affecting PSD performance. In this work, we develop a method for selecting the best performing combination of integration gates, or contiguous summed segments of the digitized pulse for PSD. The method uses a discrimination score based on the GMM PSD approach. The final selection is performed using Bayesian Optimization. PSD detection results are compared with varying numbers of selected gates on time-of-flight (TOF) data. This method can be used to fully automate the selection of gates in an unsupervised (without ground truth labels) setting.
We present a novel neural network architecture called AutoAtlas for fully unsupervised partitioning and representation learning of 3D brain Magnetic Resonance Imaging (MRI) volumes. AutoAtlas consists of two neural network components: one neural network to perform multi-label partitioning based on local texture in the volume, and a second neural network to compress the information contained within each partition. We train both of these components simultaneously by optimizing a loss function that is designed to promote accurate reconstruction of each partition, while encouraging spatially smooth and contiguous partitioning, and discouraging relatively small partitions. We show that the partitions adapt to the subject specific structural variations of brain tissue while consistently appearing at similar spatial locations across subjects. AutoAtlas also produces very low dimensional features that represent local texture of each partition. We demonstrate prediction of metadata associated with each subject using the derived feature representations and compare the results to prediction using features derived from FreeSurfer anatomical parcellation. Since our features are intrinsically linked to distinct partitions, we can then map values of interest, such as partition-specific feature importance scores onto the brain for visualization.
Analysis of longitudinal Electronic Health Record (EHR) data is an important goal for precision medicine. Difficulty in applying Machine Learning (ML) methods, either predictive or unsupervised, stems in part from the heterogeneity and irregular sampling of EHR data. We present an unsupervised probabilistic model that captures nonlinear relationships between variables over continuous-time. This method works with arbitrary sampling patterns and captures the joint probability distribution between variable measurements and the time intervals between them. Inference algorithms are derived that can be used to evaluate the likelihood of future using under a trained model. As an example, we consider data from the United States Veterans Health Administration (VHA) in the areas of diabetes and depression. Likelihood ratio maps are produced showing the likelihood of risk for moderate-severe vs minimal depression as measured by the Patient Health Questionnaire-9 (PHQ-9).
We develop an unsupervised probabilistic model for heterogeneous Electronic Health Record (EHR) data. Utilizing a mixture model formulation, our approach directly models sequences of arbitrary length, such as medications and laboratory results. This allows for subgrouping and incorporation of the dynamics underlying heterogeneous data types. The model consists of a layered set of latent variables that encode underlying structure in the data. These variables represent subject subgroups at the top layer, and unobserved states for sequences in the second layer. We train this model on episodic data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The resulting properties of the trained model generate novel insight from these complex and multifaceted data. In addition, we show how the model can be used to analyze sequences that contribute to assessment of mortality likelihood.
Prognoses of Traumatic Brain Injury (TBI) outcomes are neither easily nor accurately determined from clinical indicators. This is due in part to the heterogeneity of damage inflicted to the brain, ultimately resulting in diverse and complex outcomes. Using a data-driven approach on many distinct data elements may be necessary to describe this large set of outcomes and thereby robustly depict the nuanced differences among TBI patients' recovery. In this work, we develop a method for modeling large heterogeneous data types relevant to TBI. Our approach is geared toward the probabilistic representation of mixed continuous and discrete variables with missing values. The model is trained on a dataset encompassing a variety of data types, including demographics, blood-based biomarkers, and imaging findings. In addition, it includes a set of clinical outcome assessments at 3, 6, and 12 months post-injury. The model is used to stratify patients into distinct groups in an unsupervised learning setting. We use the model to infer outcomes using input data, and show that the collection of input data reduces uncertainty of outcomes over a baseline approach. In addition, we quantify the performance of a likelihood scoring technique that can be used to self-evaluate the extrapolation risk of prognosis on unseen patients.
Background: Functional connectomes (FCs), have been shown to provide a reproducible individual fingerprint, which has opened the possibility of personalized medicine for neuro/psychiatric disorders. Thus, developing accurate ways to compare FCs is essential to establish associations with behavior and/or cognition at the individual-level. Methods: Canonically, FCs are compared using Pearson's correlation coefficient of the entire functional connectivity profiles. Recently, it has been proposed that the use of geodesic distance is a more accurate way of comparing functional connectomes, one which reflects the underlying non-Euclidean geometry of the data. Computing geodesic distance requires FCs to be positive-definite and hence invertible matrices. As this requirement depends on the fMRI scanning length and the parcellation used, it is not always attainable and sometimes a regularization procedure is required. Results: In the present work, we show that regularization is not only an algebraic operation for making FCs invertible, but also that an optimal magnitude of regularization leads to systematically higher fingerprints. We also show evidence that optimal regularization is dataset-dependent, and varies as a function of condition, parcellation, scanning length, and the number of frames used to compute the FCs. Discussion: We demonstrate that a universally fixed regularization does not fully uncover the potential of geodesic distance on individual fingerprinting, and indeed could severely diminish it. Thus, an optimal regularization must be estimated on each dataset to uncover the most differentiable across-subject and reproducible within-subject geodesic distances between FCs. The resulting pairwise geodesic distances at the optimal regularization level constitute a very reliable quantification of differences between subjects.
Pulse shape discriminating scintillator materials in many cases allow the user to identify two basic kinds of pulses arising from two kinds of particles: neutrons and gammas, respectively. An uncomplicated solution for building a classifier consists of a two-component mixture model learned from mixtures of pulses from neutrons and gammas at a range of energies. Depending on the conditions of data gathered to be classified, multiple classes of events besides neutrons and gammas may occur, most notably pileup events. All these kinds of events that are neither neutron nor gamma are anomalous and, in cases where the class of the particle is in doubt, it is preferable to remove them from the analysis. This study compares the performance of two analytical methods for using the scores from the two-component model to identify anomalous events and in particular to remove pileup events. This study further benchmarks the analytical methods against supervised machine learning methods. This study also presents a means of assessing performance of pileup removal using ROC curves and precision–recall curves. A specific outcome of this study is to propose a novel anomaly score, denoted by G, from an unsupervised two-component model that is conveniently distributed on the interval [−1,1].
The assessment of brain fingerprints has emerged in the recent years as an important tool to study individual differences and to infer quality of neuroimaging datasets. Studies so far have mainly focused on connectivity fingerprints between different brain scans of the same individual. Here, we extend the concept of brain connectivity fingerprints beyond test/retest and assess fingerprint gradients in young adults by developing an extension of the differential identifiability framework. To do so, we look at the similarity between not only the multiple scans of an individual (subject fingerprint), but also between the scans of monozygotic and dizygotic twins (twin fingerprint). We have carried out this analysis on the 8 fMRI conditions present in the Human Connectome Project -- Young Adult dataset, which we processed into functional connectomes (FCs) and timeseries parcellated according to the Schaefer Atlas scheme, which has multiple levels of resolution. Our differential identifiability results show that the fingerprint gradients based on genetic and environmental similarities are indeed present when comparing FCs for all parcellations and fMRI conditions. Importantly, only when assessing optimally reconstructed FCs, we fully uncover fingerprints present in higher resolution atlases. We also study the effect of scanning length on subject fingerprint of resting-state FCs to analyze the effect of scanning length and parcellation. In the pursuit of open science, we have also made available the processed and parcellated FCs and timeseries for all conditions for ~1200 subjects part of the HCP-YA dataset to the scientific community.
Functional magnetic resonance imaging (fMRI) is a neuroimaging modality that captures the blood oxygen level in a subject's brain while the subject either rests or performs a variety of functional tasks under different conditions. Given fMRI data, the problem of inferring the task, known as task state decoding, is challenging due to the high dimensionality (hundreds of million sampling points per datum) and complex spatio-temporal blood flow patterns inherent in the data. In this work, we propose to tackle the fMRI task state decoding problem by casting it as a 4D spatio-temporal classification problem. We present a novel architecture called Brain Attend and Decode (BAnD), that uses residual convolutional neural networks for spatial feature extraction and self-attention mechanisms for temporal modeling. We achieve significant performance gain compared to previous works on a 7-task benchmark from the large-scale Human Connectome Project-Young Adult (HCP-YA) dataset. We also investigate the transferability of BAnD's extracted features on unseen HCP tasks, either by freezing the spatial feature extraction layers and retraining the temporal model, or finetuning the entire model. The pre-trained features from BAnD are useful on similar tasks while finetuning them yields competitive results on unseen tasks/conditions.
Microfluidic-based microencapsulation requires significant oversight to prevent material and quality loss due to sporadic disruptions in fluid flow that routinely arise. State-of-the-art microcapsule production is laborious and relies on experts to monitor the process, e.g. through a microscope. Unnoticed defects diminish the quality of collected material and/or may cause irreversible clogging. To address these issues, we developed an automated monitoring and sorting system that operates on consumer-grade hardware in real-time. Using human-labeled microscope images acquired during typical operation, we train a convolutional neural network that assesses microencapsulation. Based on output from the machine learning algorithm, an integrated valving system collects desirable microcapsules or diverts waste material accordingly. Although the system notifies operators to make necessary adjustments to restore microencapsulation, we can extend the system to automate corrections. Since microfluidic-based production platforms customarily collect image and sensor data, machine learning can help to scale up and improve microfluidic techniques beyond microencapsulation.
Current microfluidic-based microencapsulation systems rely on human experts to monitor and oversee the entire process spanning hours in order to detect and rectify when defects are found. This results in high labor costs, degradation and loss of quality in the desired collected material, and damage to the physical device. We propose an automated monitoring and classification system based on deep learning techniques to train a model for image classification into four discrete states. Then we develop an actuation control system to regulate the flow of material based on the predicted states. Experimental results of the image classification model show class average recognition rate of 95.5%. In addition, simulated test runs of our valve control system verify its robustness and accuracy.
are more easily recognized. Recent breakthroughs in Deep Learning have made it possible to learn these features from large amounts of labeled data. The focus of this project is to develop and extend Deep Learning algorithms for learning features from vast amounts of unlabeled data and to develop the HPC neural network training platform to support the training of massive network models. This LDRD project succeeded in developing new unsupervised feature learning algorithms for images and video and created a scalable neural network training toolkit for HPC. Additionally, this LDRD helped create the world's largest freely-available image and video dataset supporting open multimedia research and used this dataset for training our deep neural networks. This research helped LLNL capture several work-for-others (WFO) projects, attract new talent, and establish collaborations with leading academic and commercial partners. Finally, this project demonstrated the successful training of the largest unsupervised image neural network using HPC resources and helped establish LLNL leadership at the intersection of Machine Learning and HPC research.
Real-time emotional state recognition from neural activities has great potentials, e.g., enabling closed-loop systems to treat neuropsychiatric disorders. Besides its utility in identifying epileptogenic brain regions, electrocorticography (ECoG) provides the opportunity to study brain activities during various emotional events. In this study, we aim to test if the second order statistics such as the electrode network connectivity are relevant in discriminating different emotions. Specifically, we adopt the statistical dependence as the connectivity measure and use sparse logistic regression for classification based on the connectivity features. The practical issues of limited samples and imbalanced data are addressed. We show that the data with the highest synchronization in the γ band provide the most competitive performance. The identified brain regions involved in the most discriminative connectivity patterns are consistent with the previous findings in the literature.
Joseph a Osullivan合作论文数Electrical and Systems Engineering Department7