There remains an open question about the usefulness and the interpretation of machine learning (ML) approaches for discrimination of spatial patterns of brain images between samples or activation states. In the last few decades, these approaches have limited their operation to feature extraction and linear classification tasks for between-group inference. In this context, statistical inference is assessed by randomly permuting image labels or by the use of random effect models that consider between-subject variability. These multivariate ML-based statistical pipelines, whilst potentially more effective for detecting activations than hypotheses-driven methods, have lost their mathematical elegance, ease of interpretation, and spatial localization of the ubiquitous General linear Model (GLM). Recently, the estimation of the conventional GLM parameters has been demonstrated to be connected to an univariate classification task when the design matrix in the GLM is expressed as a binary indicator matrix. In this paper we explore the complete connection between the univariate GLM and ML-based regressions. To this purpose we derive a refined statistical test with the GLM based on the parameters obtained by a linear Support Vector Regression (SVR) in the inverse problem (SVR-iGLM). Subsequently, random field theory (RFT) is employed for assessing statistical significance following a conventional GLM benchmark. Experimental results demonstrate how parameter estimations derived from each model (mainly GLM and SVR) result in different experimental design estimates that are significantly related to the predefined functional task. Moreover, using real data from a multisite initiative the proposed ML-based inference demonstrates statistical power and the control of false positives, outperforming the regular GLM.
In the 70s a novel branch of statistics emerged focusing its effort in selecting a function in the pattern recognition problem, which fulfils a definite relationship between the quality of the approximation and its complexity. These data-driven approaches are mainly devoted to problems of estimating dependencies with limited sample sizes and comprise all the empirical out-of sample generalization approaches, e.g. cross validation (CV) approaches. Although the latter are \emph{not designed for testing competing hypothesis or comparing different models} in neuroimaging, there are a number of theoretical developments within this theory which could be employed to derive a Statistical Agnostic (non-parametric) Mapping (SAM) at voxel or multi-voxel level. Moreover, SAMs could relieve i) the problem of instability in limited sample sizes when estimating the actual risk via the CV approaches, e.g. large error bars, and provide ii) an alternative way of Family-wise-error (FWE) corrected p-value maps in inferential statistics for hypothesis testing. In this sense, we propose a novel framework in neuroimaging based on concentration inequalities, which results in (i) a rigorous development for model validation with a small sample/dimension ratio, and (ii) a less-conservative procedure than FWE p-value correction, to determine the brain significance maps from the inferences made using small upper bounds of the actual risk.
In the immediate future, with the increasing presence of electrical vehicles and the large increase in the use of renewable energies, it will be crucial that distribution power networks are managed, supervised and exploited in a similar way as the transmission power systems were in previous decades. To achieve this, the underlying infrastructure requires automated monitoring and digitization, including smart-meters, wide-band communication systems, electronic device based-local controllers, and the Internet of Things. All of these technologies demand a huge amount of data to be curated, processed, interpreted and fused with the aim of real-time predictive control and supervision of medium/low voltage transformer substations. Wiener–Granger causality, a statistical notion of causal inference based on Information Fusion could help in the prediction of electrical behaviour arising from common causal dependencies. Originally developed in econometrics, it has successfully been applied to several fields of research such as the neurosciences and is applicable to time series data whereby cause precedes effect. In this paper, we demonstrate the potential of this methodology in the context of power measures for providing theoretical models of low/medium power transformers. Up to our knowledge, the proposed method in this context is the first attempt to build a data-driven power system model based on G-causality. In particular, we analysed directed functional connectivity of electrical measures providing a statistical description of observed responses, and identified the causal structure within data in an exploratory analysis. Pair-wise conditional G-causality of power transformers, their independent evolution in time, and the joint evolution in time and frequency are discussed and analysed in the experimental section.
This paper deals with the topic of learning from unlabeled or noisy-labeled data in the context of a classification problem. In the classification problem the outcome yields one of a discrete set of values thus, assumptions on them could be established to obtain the most likely prediction model at the training stage. In this paper, a novel case-based model selection method is proposed, which combines hypothesis testing from a discrete set of expected outcomes and feature extraction within a cross-validated classification stage. This wrapper-type procedure acts on fully-observable variables under hypothesis-testing and improves the classification accuracy on the test set, or keeps its performance at least at the level of the statistical classifier. The model selection strategy in the cross validation loop allows building an ensemble classifier that could improve the performance of any expert and intelligence system, particularly on small sample-size datasets. Experiments were carried out on several databases yielding a clear improvement on the baseline, i.e., SPECT dataset Acc = 86.35 +/- 1.51, with Sen = 91.10 +/- 2.77, and Spe = 81.11 +/- 1.61. In addition, the CV error estimate for the classifier under our approach was found to be an almost unbiased estimate (as the baseline approach) of the true error that the classifier would incur on independent data. (C) 2017 Elsevier Ltd. All rights reserved.
Background: Alzheimer's disease (AD) is the most common cause of dementia in the elderly and affects approximately 30 million individuals worldwide. Mild cognitive impairment (MCI) is very frequently a prodromal phase of AD, and existing studies have suggested that people with MCI tend to progress to AD at a rate of about 10-15% per year. However, the ability of clinicians and machine learning systems to predict AD based on MRI biomarkers at an early stage is still a challenging problem that can have a great impact in improving treatments. Method: The proposed system, developed by the SiPBA-UGR team for this challenge, is based on feature standardization, ANOVA feature selection, partial least squares feature dimension reduction and an ensemble of One vs. Rest random forest classifiers. With the aim of improving its performance when discriminating healthy controls (HC) from MCI, a second binary classification level was introduced that reconsiders the HC and MCI predictions of the first level. Results: The system was trained and evaluated on an ADNI datasets that consist of T1-weighted MRI morphological measurements from HC, stable MCI, converter MCI and AD subjects. The proposed system yields a 56.25% classification score on the test subset which consists of 160 real subjects. Comparison with existing method(s): The classifier yielded the best performance when compared to: (i) One vs. One (OvO), One vs. Rest (OvR) and error correcting output codes (ECOC) as strategies for reducing the multiclass classification task to multiple binary classification problems, (ii) support vector machines, gradient boosting classifier and random forest as base binary classifiers, and (iii) bagging ensemble learning. Conclusions: A robust method has been proposed for the international challenge on MCI prediction based on MRI data. The system yielded the second best performance during the competition with an accuracy rate of 56.25% when evaluated on the real subjects of the test set. (C) 2017 Elsevier B.V. All rights reserved.
Nowadays, 35 million people worldwide suffer from some form of dementia. Given the increase in life expectancy it is estimated that in 2035 this number will grow to 115 million. Alzheimer's disease is the most common cause of dementia and it is of great importance diagnose it at an early stage. This is the main goal of this work, the development of a new automatic method to predict the mild cognitive impairment (MCI) patients who will develop Alzheimer's disease within one year or, conversely, its impairment will remain stable. This technique will analyze data from both magnetic resonance imaging and neuropsychological tests by utilizing a t-test for feature selection, maximum-uncertainty linear discriminant analysis (MLDA) for classification and leave-one-out cross validation (LOOCV) for evaluating the performance of the methods, which achieved a classification accuracy of 73.95 %, with a sensitivity of 72.14 % and a specificity of 73.77 %.
Alzheimer's disease (AD) is the most common cause of dementia. Nowadays, 44 million people worldwide suffer from this neurodegenerative disease. Fortunately, the use of new technologies can help doctors in diagnosing this disease in an increasingly early stage, which is vital to prevent its advance. In this work we have developed a new automatic method to predict if patients suffering from mild cognitive impairment (MCI) will develop AD within one year or, conversely, its impairment will remain stable. This technique is based on the so-called Searchlight, a widely known approach in fMRI but which has not been previously used with structural images. Besides analyzing the intensity of the voxels in each of the subregions defined by the Searchlight, data from two neuro-psychological tests were used during the classification process, achieving an accuracy of 84%.
Positron emission tomography (PET) provides a functional imaging modality to detect signs of dementias in human brains. Two-dimensional empirical mode decomposition (2D-EMD) provides means to analyze such images. It decomposes the latter into characteristic modes which represent textures on different spatial scales. These textures provide informative features for subsequent classification purposes. The study proposes a new EMD variant which relies on a Green's function based estimation method including a tension parameter to fast and reliably estimate the envelope hypersurfaces interpolating extremal points of the two-dimensional intensity distrubution of the images. The new method represents a fast and stable bi-dimensional EMD which speeds up computations roughly 100-fold. In combination with proper classifiers these exploratory feature extraction techniques can form a computer aided diagnosis (CAD) system to assist clinicians in identifying various diseases from functional images alone. PET images of subjects suffering from Alzheimer's disease are taken to illustrate this ability.
An analysis of binary data sets employing Bernoulli statistics and a partially non-negative factorization of the related matrix of log-odds is presented. The model places several constraints onto the factorization process rendering the estimated basis system strictly non-negative or even binary. Thereby the proposed model places itself in between a logistic PCA and a binary NMF approach. We show with proper toy data sets that different variants of the proposed model yield reasonable results and indeed are able to estimate with good precision the underlying basis system which forms a new and often more compact representation of the observations. An application of the method to the USPS data set reveals the performance of the various variants of the model and shows good reconstruction quality even with a low rank binary basis set.
Every year, malaria kills between 660,000 and 1.2 million people, many of whom are children in Africa. The World Health Organization (WHO) encourages the development of rapid and economical diagnostic tests that allow for the identification of proper treatment methods. In this paper a novel method to automatically enumerate malaria parasites is proposed and evaluated, using a database consisting of 475 images with varying densities of malaria parasites. This method will analyze data by utilizing standard operations of image processing such as histogram equalization, thresholding, morphological operations and connected components analysis for parasite density estimation. The application of the proposed method yields an average accuracy rate of 96.46% with a low processing time of two seconds per image on a custom computing platform. (C) 2014 Elsevier Ltd. All rights reserved.
This paper shows an adaptive statistical test for QRS detection of electrocardiography (ECG) signals. The method is based on a M-ary generalized likelihood ratio test (LRT) defined over a multiple observation window in the Fourier domain. The motivations for proposing another detection algorithm based on maximum a posteriori (MAP) estimation are found in the high complexity of the signal model proposed in previous approaches which i) makes them computationally unfeasible or not intended for real time applications such as intensive care monitoring and (ii) in which the parameter selection conditions the overall performance. In this sense, we propose an alternative model based on the independent Gaussian properties of the Discrete Fourier Transform (DFT) coefficients, which allows to define a simplified MAP probability function. In addition, the proposed approach defines an adaptive MAP statistical test in which a global hypothesis is defined on particular hypotheses of the multiple observation window. In this sense, the observation interval is modeled as a discontinuous transmission discrete-time stochastic process avoiding the inclusion of parameters that constraint the morphology of the QRS complexes.
In last years, many research eorts in neurosciences have focused in multivariate approaches based on machine learning as an al- ternative to the use of Statistical Parametric Mapping and the univariate analyses that it provides. However, this relatively new eld still lacks of a software framework that completely meets the needs of the scientic community. In this work we present a toolbox designed to facilitate the access to the recent advances in neuroimaging data analysis based on multivariate approaches. The toolbox, written on Matlab, is freely avail- able and implements a Graphical User Interface that allows managing neuroimaging data in an easy way.
Positron emission tomography (PET) provides a functional imaging modality to detect signs of dementias in human brains. Two-dimensional empirical mode decomposition (2D EMD) provides means to analyze such images. It extracts characteristic textures from these images which may be fed into powerful classifiers trained to group these textures into several classes depending on the problem at hand. The study investigates the potential use of 2D EEMD in combination with proper classifiers to form a computer aided diagnosis (CAD) system to assist clinicians in identifying various diseases from functional images alone. PET images of subjects suffering from a dementia are taken to illustrate this ability.
The existence of an endophenotype of autism spectrum condition (ASC) has been recently suggested by several commentators. It can be estimated by finding differences between controls and people with ASC that are also present when comparing controls and the unaffected siblings of ASC individuals. In this work, we used a multivariate methodology applied on magnetic resonance images to look for such differences. The proposed procedure consists of combining a searchlight approach and a support vector machine classifier to identify the differences between three groups of participants in pairwise comparisons: controls, people with ASC and their unaffected siblings. Then we compared those differences selecting spatially collocated as candidate endophenotypes of ASC.
This paper shows an adaptive statistical test for QRS detection of ECG signals. The method is based on a M-ary generalized likelihood ratio test (LRT) defined over a multiple observation window in the Fourier domain. The previous algorithms based on maximum a posteriori (MAP) estimation result in high signal model complexity which i) makes them computationally unfeasible or not intended for real time applications such as intensive care monitoring and (ii) in which the parameter selection conditions the overall performance. A simplified model based on the independent Gaussian properties of the DFT coefficients is proposed. This model allows to define a simplified MAP probability function and to define an adaptive MAP statistical test in which a global hypothesis is defined on particular hypotheses of the multiple observation window. Moreover, the observation interval is modeled as a discontinuous transmission discrete-time stochastic process avoiding the inclusion of parameters that constraint the morphology of the QRS complexes.
NMF is a blind source separation technique decomposing multivariate non-negative data sets into meaningful non-negative basis components and non-negative weights. There are still open problems to be solved: uniqueness and model order selection as well as developing efficient NMF algorithms for large scale problems. Addressing uniqueness issues, we propose a Bayesian optimality criterion (BOC) for NMF solutions which can be derived in the absence of prior knowledge. Furthermore, we present a new Variational Bayes NMF algorithm VBNMF which is a straight forward generalization of the canonical Lee–Seung method for the Euclidean NMF problem and demonstrate its ability to automatically detect the actual number of components in non-negative data.
A novel Score-based Physarum Learner algorithm for learning Bayesian Network structure from data is introduced and shown to outperform common score based structure learning algorithms for some benchmark data sets. The Score-based Physarum Learner first initializes a fully connected Physarum-Maze with random conductances. In each Physarum Solver iteration, the source and sink nodes are changed randomly, and the conductances are updated. Connections exceeding a predefined conductance threshold are considered as Bayesian Network edges, and the score of the connected nodes are examined in both directions. A positive or negative feedback is given to the edge conductance based on the calculated scores. Due to randomness in selecting connections for evaluation, an ensemble of Score-based Physarum Learner is used to build the final Bayesian Network structure. (c) 2014 Elsevier Ltd. All rights reserved.
The use of functional imaging has been proven very helpful for the process of diagnosis of neurodegenerative diseases, such as Alzheimer's Disease (AD). In many cases, the analysis of these images is performed by manual reorientation and visual interpretation. Therefore, new statistical techniques to perform a more quantitative analysis are needed. In this work, a new statistical approximation to the analysis of functional images, based on significance measures and Independent Component Analysis (ICA) is presented. After the images preprocessing, voxels that allow better separation of the two classes are extracted, using significance measures such as the Mann-Whitney-Wilcoxon U-Test (MWW) and Relative Entropy (RE). After this feature selection step, the voxels vector is modelled by means of ICA, extracting a few independent components which will be used as an input to the classifier. Naive Bayes and Support Vector Machine (SVM) classifiers are used in this work. The proposed system has been applied to two different databases. A 96-subjects Single Photon Emission Computed Tomography (SPECT) database from the "Virgen de las Nieves" Hospital in Granada, Spain, and a 196-subjects Positron Emission Tomography (PET) database from the Alzheimer's Disease Neuroimaging Initiative (ADNI). Values of accuracy up to 96.9% and 91.3% for SPECT and PET databases are achieved by the proposed system, which has yielded many benefits over methods proposed on recent works.
In this work, a novel approach to Computer Aided Diagnosis (CAD) system for the Parkinson’s Disease (PD) is proposed. This tool is intended for physicians, and is based on fully automated methods that lead to the classification of Ioflupane/FP-CIT-I-123 (DaTSCAN) SPECT images. DaTSCAN images from the Parkinson Progression Markers Initiative (PPMI) are used to have in vivo information of the dopamine transporter density. These images are normalized, reduced (using a mask), and then a GLC matrix is computed over the whole image, extracting several Haralick texture features which will be used as a feature vector in the classification task. Using the leave-one-out cross-validation technique over the whole PPMI database, the system achieves results up to a 95.9% of accuracy, and 97.3% of sensitivity, with positive likelihood ratios over 19, demonstrating our system’s ability on the detection of the Parkinson’s Disease by providing robust and accurate results for clinical practical use, as well as being fast and automatic.
An automated method for orientation of functional brain image is proposed. Intrinsec information is captured from the image in three stages: first the volume to identify the anterior to posterior line, second the symmetry to detect the hemisphere dividing plane and third the contour to determine the up-down and front-back orientation. The approach is tested in more than a tousand images from different formats and modalities with high reconition rates.