BACKGROUND: Autism spectrum disorder (ASD) is among the most pervasive neurodevelopmental disorders, yet the neurobiology of ASD is still poorly understood because inconsistent findings from underpowered individual studies preclude the identification of robust and interpretable neurobiological markers and predictors of clinical symptoms. METHODS: We leverage multiple brain imaging cohorts and exciting recent advances in explainable artificial intelligence to develop a novel spatiotemporal deep neural network (stDNN) model, which identifies robust and interpretable dynamic brain markers that distinguish ASD from neurotypical control subjects and predict clinical symptom severity. RESULTS: stDNN achieved consistently high classification accuracies in cross-validation analysis of data from the multisite ABIDE (Autism Brain Imaging Data Exchange) cohort (n = 834). Crucially, stDNN also accurately classified data from independent Stanford ( n = 202) and GENDAAR (Gender Exploration of Neurogenetics and Development to Advanced Autism Research) (n = 90) cohorts without additional training. stDNN could not distinguish attentiondeficit/hyperactivity disorder from neurotypical control subjects, highlighting the model's specificity. Explainable artificial intelligence revealed that brain features associated with the posterior cingulate cortex and precuneus, dorsolateral and ventrolateral prefrontal cortex, and superior temporal sulcus, which anchor the default mode network, cognitive control, and human voice processing systems, respectively, most clearly distinguished ASD from neurotypical control subjects in the three cohorts. Furthermore, features associated with the posterior cingulate cortex and precuneus nodes of the default mode network emerged as robust predictors of the severity of core social and communication deficits but not restricted/repetitive behaviors in ASD. CONCLUSIONS: Our findings, replicated across independent cohorts, reveal robust individualized functional brain fingerprints of ASD psychopathology, which could lead to more objective and precise phenotypic characterization and targeted treatments.
Sea ice observations through satellite imaging have led to advancements in environmental research, ship navigation, and ice hazard forecasting in cold regions. Machine learning and, recently, deep learning techniques are being explored by various researchers to process vast amounts of Synthetic Aperture Radar (SAR) data for detecting potential hazards in navigational routes. Detection of hazards such as sea ice floes in Marginal Ice Zones (MIZs) is quite challenging as the floes are often embedded in a multiscale ice cover composed of ice filaments and eddies in addition to floes. This study proposes a segmentation model tailored for detecting ice floes in SAR images. The model exploits the advantages of both convolutional neural networks and convolutional conditional random field (Conv-CRF) in a combined manner. The residual UNET (RES-UNET) computes expressive features to generate coarse segmentation maps while the Conv-CRF exploits the spatial co-occurrence pairwise potentials along with the RES-UNET unary/segmentation maps to generate final predictions. The whole pipeline is trained end-to-end using a dual loss function. This dual loss function is composed of a weighted average of binary cross entropy and soft dice loss. The comparison of experimental results with the conventional segmentation networks such as UNET, DeepLabV3, and FCN-8 demonstrates the effectiveness of the proposed architecture.
Introduction Autism spectrum disorder (ASD) is among the most common and pervasive neurodevelopmental disorders. Yet, despite decades of research, the neurobiology of ASD is still poorly understood, as inconsistent findings preclude the identification of robust and interpretable neurobiological markers and predictors of clinical symptoms. Objectives Identify robust and interpretable dynamic brain markers that distinguish children with ASD from typically-developing (TD) children and predict clinical symptom severity. Methods We leverage multiple functional brain imaging cohorts (ABIDE, Stanford; N = 1004) and exciting recent advances in explainable artificial intelligence (xAI), to develop a novel multivariate time series deep neural network model that extracts informative brain dynamics features that accurately distinguish between ASD and TD children, and predict clinical symptom severity. Results Our model achieved consistently high classification accuracies in cross-validation analysis of data from the ABIDE cohort. Crucially, despite the differences in symptom profiles, age, and data acquisition protocols, our model also accurately classified data from an independent Stanford cohort without additional training. xAI analyses revealed that brain features associated with the default mode network, and the human voice/face processing and communication systems, most clearly distinguished ASD from TD children in both cohorts. Furthermore, the posterior cingulate cortex emerged as robust predictor of the severity of social and communication deficits in ASD in both cohorts. Conclusions Our findings, replicated across two independent cohorts, reveal robust and neurobiologically interpretable brain features that detect ASD and predict core phenotypic features of ASD, and have the potential to transform our understanding of the etiology and treatment of the disorder. Disclosure No significant relationships.
An ongoing major challenge in computer vision is the task of person re-identification, where the goal is to match individuals across different, non-overlapping camera views. While recent success has been achieved via supervised learning using deep neural networks, such methods have limited widespread adoption due to the need for large-scale, customized data annotation. As such, there has been a recent focus on unsupervised learning approaches to mitigate the data annotation issue; however, current approaches in literature have limited performance compared to supervised learning approaches as well as limited applicability for adoption in new environments. In this paper, we address the aforementioned challenges faced in person re-identification for real-world, practical scenarios by introducing a novel, unsupervised domain adaptation approach for person re-identification. This is accomplished through the introduction of: i) k-reciprocal tracklet Clustering for Unsupervised Domain Adaptation (ktCUDA) (for pseudo-label generation on target domain), and ii) Synthesized Heterogeneous RE-id Domain (SHRED) composed of large-scale heterogeneous independent source environments (for improving robustness and adaptability to a wide diversity of target environments). Experimental results across four different image and video benchmark datasets show that the proposed ktCUDA and SHRED approach achieves an average improvement of +5.7 mAP in re-identification performance when compared to existing state-of-the-art methods, as well as demonstrate better adaptability to different types of environments.
In this study, we propose the leveraging of interpretability for tasks beyond purely the purpose of explainability. In particular, this study puts forward a novel strategy for leveraging gradient-based interpretability in the realm of adversarial examples, where we use insights gained to aid adversarial learning. More specifically, we introduce the concept of spatially constrained one-pixel adversarial perturbations, where we guide the learning of such adversarial perturbations towards more susceptible areas identified via gradient-based interpretability. Experimental results using different benchmark datasets show that such a spatially constrained one-pixel adversarial perturbation strategy can noticeably improve the speed of convergence as well as produce successful attacks that were also visually difficult to perceive, thus illustrating an effective use of interpretability methods for tasks outside of the purpose of purely explainability.
Soil organic carbon (SOC) content is key component of the global carbon (C) cycle which is highly variable with respect to space and time. The main objective of this study was to provide an assessment of soil organic carbon (SOC) stock variability for Uttarakhand state. The other objective of this study was to evaluate the performance of different pedotransfer functions for reliable assessment of bulk density. Soil Resource Mapping for Uttarakhand state was conducted on 1:50,000 scale with the help of Satellite imagery (LISS III) along with exhaustive ground truthing through soil surveys. Stratified sampling was carried out based on remotely sensed satellite data for different slope, physiography and land-use/cover. The physico-chemical properties of selected samples for agriculture and forest land use were utilized for analyzing the performance of six pedotransfer functions for assessment of bulk density. The SOC stocks were estimated on the basis of soil organic matter content for top 20 cm layer and bulk density estimated from best performing pedotransfer functions models. The SOC stock class of 51-100 tonnes C ha-1 was dominated by covering 42.00% of state area followed by 26-50 tonnes C ha-1 class covering 23.74% area. Similarly, about 7.91% and 3.24 % area of state are covered under 11-25 tonnes C ha-1 and 101-160 tonnes C ha-1 classes, respectively. Remaining 22.44 % of state not forms part of study were mapped under settlement, snowbound area, drainages/rivers, reservoirs etc. The difference in performance of pedotransfer functions under different land use system implies the necessity of evaluation of pedotransfer functions before their implementation. Significantly greater SOC stocks were observed in forest and grassland/open-scrub land use and such differences can be attributed to the higher tree/shrub density, shrub/herb biomass and forest litter in the forest areas as compared to agriculture land use.
Radiomics-driven computer aided diagnosis (CAD) has shown considerable promise in recent years as a potential tool for improving clinical decision support in medical oncology, particularly those based around the concept of discovery radiomics, where radiomic sequencers are discovered through the analysis of medical imaging data. One of the main limitations, with current CAD approaches, is that it is very difficult to gain insight or rationale as to how decisions are made, thus limiting their utility to clinicians. In this paper, we propose CLEAR-DR, a novel interpretable CAD system based on the notion of CLass-Enhanced Attentive Response Discovery Radiomics for the purpose of clinical decision support for diabetic retinopathy. In addition to disease grading via the discovered deep radiomic sequencer, the CLEAR-DR system also produces a visual interpretation of the decision-making process to provide better insight and understanding of the decision-making process of the system. We demonstrate the effectiveness and utility of the proposed CLEAR-DR system of enhancing the interpretability of diagnostic grading results for the application of diabetic retinopathy grading. CLEAR-DR can act as a potentially powerful tool to address the uninterpretability issue of current CAD systems, thus improving their utility to clinicians.
Lung cancer is the leading cause of cancer-related death worldwide. Computer-aided diagnosis (CAD) systems have shown significant promise in recent years for facilitating the effective detection and classification of abnormal lung nodules in computed tomography (CT) scans. While hand-engineered radiomic features have been traditionally used for lung cancer prediction, there have been significant recent successes achieving state-of-the-art results in the area of discovery radiomics. Here, radiomic sequencers comprising of highly discriminative radiomic features are discovered directly from archival medical data. However, the interpretation of predictions made using such radiomic sequencers remains a challenge. A novel end-to-end interpretable discovery radiomics-driven lung cancer prediction pipeline has been designed, build, and tested. The radiomic sequencer being discovered possesses a deep architecture comprised of stacked interpretable sequencing cells (SISC). The SISC architecture is shown to outperform previous approaches while providing more insight in to its decision making process. The SISC radiomic sequencer is able to achieve state-of-the-art results in lung cancer prediction, and also offers prediction interpretability in the form of critical response maps. The critical response maps are useful for not only validating the predictions of the proposed SISC radiomic sequencer, but also provide improved radiologist-machine collaboration for effective diagnosis.
Person re-identification (ReID) remains a very difficult challenge in computer vision, and critical for large-scale video surveillance scenarios where an individual could appear in different camera views at different times. There has been recent interest in tackling this challenge using cross-domain approaches, which leverages data from source domains that are different than the target domain. Such approaches are more practical for real-world widespread deployment given that they don't require on-site training (as with unsupervised or domain transfer approaches) or on-site manual annotation and training (as with supervised approaches). In this study, we take a systematic approach to establishing a large baseline source domain and target domain for cross-domain person ReID. We accomplish this by conducting a comprehensive analysis to study the similarities between source domains proposed in literature, and studying the effects of incrementally increasing the size of the source domain. This allows us to establish a balanced source domain and target domain split that promotes variety in both source and target domains. Furthermore, using lessons learned from the state-of-the-art supervised person re-identification methods, we establish a strong baseline method for cross-domain person ReID. Experiments show that a source domain composed of two of the largest person ReID domains (SYSU and MSMT) performs well across six commonly-used target domains. Furthermore, we show that, surprisingly, two of the recent commonly-used domains (PRID and GRID) have too few query images to provide meaningful insights. As such, based on our findings, we propose the following balanced baseline for cross-domain person ReID consisting of: i) a fixed multi-source domain consisting of SYSU, MSMT, Airport and 3DPeS, and ii) a multi-target domain consisting of Market-1501, DukeMTMC-reID, CUHK03, PRID, GRID and VIPeR.
One of the main challenges for broad adoption of deep learning based models such as convolutional neural networks (CNN), is the lack of understanding of their decisions. In many applications, a simpler, less capable model that can be easily understood is favorable to a black-box model that has superior performance. In this paper, we present an approach for designing CNNs based on visualization of the internal activations of the model. We visualize the model's response through attentive response maps obtained using a fractional stride convolution technique and compare the results with known imaging landmarks from the medical literature. We show that sufficiently deep and capable models can be successfully trained to use the same medical landmarks a human expert would use. Our approach allows for communicating the model decision process well, but also offers insight towards detecting biases.
Computational methods that automatically extract knowledge from data are critical for enabling data-driven materials science. A reliable identification of lattice symmetry is a crucial first step for materials characterization and analytics. Current methods require a user-specified threshold, and are unable to detect average symmetries for defective structures. Here, we propose a machine-learning-based approach to automatically classify structures by crystal symmetry. First, we represent crystals by calculating a diffraction image, then construct a deep-learning neural-network model for classification. Our approach is able to correctly classify a dataset comprising more than 100 000 simulated crystal structures, including heavily defective ones. The internal operations of the neural network are unraveled through attentive response maps, demonstrating that it uses the same landmarks a materials scientist would use, although never explicitly instructed to do so. Our study paves the way for crystal-structure recognition of - possibly noisy and incomplete - three-dimensional structural data in big-data materials science.
Robust place recognition systems are essential for long term localization and autonomy. Such systems should recognize scenes with both conditional and viewpoint changes. In this paper, we present a deep learning based planar omni-directional place recognition approach that can simultaneously cope with conditional and viewpoint variations, including large viewpoint changes, which current methods do not address. We evaluate the proposed method on two real world datasets dealing with illumination, seasonal/weather changes and changes occurred in the environment across a period of 1 year, respectively. We provide both quantitative (recall at 100% precision) and qualitative (confusion matrices) comparison of the basic pipeline of place recognition for the omni-directional approach with single-view and side-view camera approaches. The results prove the efficacy of the proposed omnidirectional deep learning method over the single-view and side-view cameras in dealing with both conditional and large viewpoint changes.
In this work, we propose CLass-Enhanced Attentive Response (CLEAR): an approach to visualize and understand the decisions made by deep neural networks (DNNs) given a specific input. CLEAR facilitates the visualization of attentive regions and levels of interest of DNNs during the decision-making process. It also enables the visualization of the most dominant classes associated with these attentive regions of interest. As such, CLEAR can mitigate some of the shortcomings of heatmap-based methods associated with decision ambiguity, and allows for better insights into the decision-making process of DNNs. Quantitative and qualitative experiments across three different datasets demonstrate the efficacy of CLEAR for gaining a better understanding of the inner workings of DNNs during the decision-making process.
Measuring nutritional intake is a tool that is critical to themonitoring of health, both as an individual or of a group. It isespecially important in the monitoring of those at risk formalnutrition, an issue which costs billions of dollars globally, andcurrent methods used in practice are manual, time-consuming,and have inherent biases and inaccuracies. This study proposes anovel imaging system with a superpixel-based segmentationalgorithm as part of an automated nutritional intake system. Thestudy also examines three important parameters of the algorithmand their ideal values; region size and spatial regularization forsuperpixel segmentation, as well as spatial weighting inclustering. The experimental results demonstrate that theproposed system is effective in segmenting an image of a plate intoits constituent foods.
Deep learning has been shown to outperform traditional machinelearning algorithms across a wide range of problem domains. However,current deep learning algorithms have been criticized as uninterpretable"black-boxes" which cannot explain their decision makingprocesses. This is a major shortcoming that prevents the widespreadapplication of deep learning to domains with regulatoryprocesses such as finance. As such, industries such as financehave to rely on traditional models like decision trees that are muchmore interpretable but less effective than deep learning for complexproblems. In this paper, we propose CLEAR-Trade, a novelfinancial AI visualization framework for deep learning-driven stockmarket prediction that mitigates the interpretability issue of deeplearning methods. In particular, CLEAR-Trade provides a effectiveway to visualize and explain decisions made by deep stock marketprediction models. We show the efficacy of CLEAR-Trade in enhancingthe interpretability of stock market prediction by conductingexperiments based on S&P 500 stock index prediction. The resultsdemonstrate that CLEAR-Trade can provide significant insightinto the decision-making process of deep learning-driven financialmodels, particularly for regulatory processes, thus improving theirpotential uptake in the financial industry.
Lung cancer is the leading cause for cancer related deaths. As such, there is an urgent need for a streamlined process that can allow radiologists to provide diagnosis with greater efficiency and accuracy. A powerful tool to do this is radiomics: a high-dimension imaging feature set. In this study, we take the idea of radiomics one step further by introducing the concept of discovery radiomics for lung cancer prediction using CT imaging data. In this study, we realize these custom radiomic sequencers as deep convolutional sequencers using a deep convolutional neural network learning architecture. To illustrate the prognostic power and effectiveness of the radiomic sequences produced by the discovered sequencer, we perform cancer prediction between malignant and benign lesions from 97 patients using the pathologically-proven diagnostic data from the LIDC-IDRI dataset. Using the clinically provided pathologically-proven data as ground truth, the proposed framework provided an average accuracy of 77.52% via 10-fold cross-validation with a sensitivity of 79.06% and specificity of 76.11%, surpassing the state-of the art method.
One of the main challenges for broad adoption of deep convolutional neural network (DCNN) models is the lack of understanding of their decision process. In many applications a simpler less capable model that can be easily understood is favorable to a black-box model that has superior performance. In this paper, we present an approach for designing DCNN models based on visualization of the internal activations of the model. We visualize the model's response using fractional stride convolution technique and compare the results with known imaging landmarks from the medical literature. We show that sufficiently deep and capable models can be successfully trained to use the same medical landmarks a human expert would use. The presented approach allows for communicating the model decision process well, but also offers insight towards detecting biases.
Robot deployment in open snow-covered environments poses challenges to existing vision-based localization and mapping methods. Limited field of view and over-exposure in regions where snow is present leads to difficulty identifying and tracking features in the environment. The wide variation in scene depth and relative visual saliency of points on the horizon results in clustered features with poor depth estimates, as well as the failure of typical keyframe selection metrics to produce reliable bundle adjustment results. In this work, we propose the use of and two extensions to Multi-Camera Parallel Tracking and Mapping (MCPTAM) to improve localization performance in snow-laden environments. First, we define a snowsegmentation method and snow-specific image filtering to enhance detectability of local features on the snow surface. Then, we define a feature entropy reduction metric for keyframe selection that leads to reduced map sizes while maintaining localization accuracy. Both refinements are demonstrated on a snow-laden outdoor dataset collected with a wide field-of-view, three camera cluster on a ground rover platform.
Radiomics has proven to be a powerful prognostic tool for cancer detection, and has previously been applied in lung, breast, prostate, and head-and-neck cancer studies with great success. However, these radiomics-driven methods rely on pre-defined, hand-crafted radiomic feature sets that can limit their ability to characterize unique cancer traits. In this study, we introduce a novel discovery radiomics framework where we directly discover custom radiomic features from the wealth of available medical imaging data. In particular, we leverage novel StochasticNet radiomic sequencers for extracting custom radiomic features tailored for characterizing unique cancer tissue phenotype. Using StochasticNet radiomic sequencers discovered using a wealth of lung CT data, we perform binary classification on 42,340 lung lesions obtained from the CT scans of 93 patients in the LIDC-IDRI dataset. Preliminary results show significant improvement over previous state-of-the-art methods, indicating the potential of the proposed discovery radiomics framework for improving cancer screening and diagnosis.
In this paper, we describe the underlying methodology behind discovery radiomics, where the ultimate goal is to discover customized, abstract radiomic feature models directly from the wealth of medical imaging data to better capture highly unique tumor traits beyond what can be captured using hand-crafted radiomic feature models. We further explore the current state-of-the-art in discovery radiomics and their application to various forms of cancer such as prostate cancer and lung cancer, and show that discovery radiomics can yield significant potential clinical impact.