Machine learning algorithms have great potential for classifying brain activity, and lightweight classifier algorithms, requiring little computational resources, can be used on low-energy neuromorphic hardware designed for implantable neuroprosthetics. One of these efficient algorithms, the Liquid State Machine, implements the concept of Spiking Neural Networks and has been shown to achieve outstanding results on the task of whisker stimulus detection from the mouse barrel cortex, a widely used model system. While this is promising for neuroprosthetics, it has been unclear how a Spiking Neural Network or other machine learning algorithms perform on data recorded from awake mice and how trained models generalize across individuals, the latter being relevant to transferring trained models to new hardware. Using laminar multi-electrode local field potential recordings obtained from four mice performing a single-whisker detection task, we benchmarked the performance of a collection of lightweight classification algorithms. We found that the Liquid State Machine, a generalized linear model, and the time series classifier ROCKET are the most accurate for stimulus detection. Among those, the Liquid State Machine achieved the fastest model training and inference runtime and provided robust accuracy across individual mice. Additional analyses show that there is no significant improvement in using multiple cortical layers as input for the model and that 40 ms of stimulus recording is sufficient to maintain high detection accuracy.
The solution of high-dimensional nonlinear regression problems through standard machine learning approaches often relies on first-order information, due to the numerical and memory challenges arising from the computation of the Hessian matrix and of the higher-order derivatives. While this scenario seems not favorable to second-order methods, here we show that an efficient and modular structure-exploiting interior-point solver can be successfully applied to the recently introduced class of entropy-based methods for regression learning. Specifically, by exploiting the favorable structure of the problem and of the Hessian matrix, we suggest a robust solution strategy based on explicit low-rank updates combined with an iterative Symmetric Quasi-Minimal Residual (SQMR) algorithm to solve the underlying system of linear equations. The results show that the proposed structure-exploiting solver—which relies on the hybrid parallelism and distributed-memory computing paradigm—allows a significant solution time speed-up with respect to a naive solution strategy. Furthermore, through an adequate use of the Message Passing Interface (MPI) and of Open Multi-Processing (OpenMP), the proposed solver enables the solution of large-scale problems on high-performance computing architectures consisting of thousands of compute nodes. The accompanying detailed convergence and performance analyses demonstrate both numerical robustness and high-performance capabilities for increasingly high-dimensional problems.
Small data learning problems are characterized by a significant discrepancy between the limited number of response variable observations and the large feature space dimension. In this setting, the common learning tools struggle to identify the features important for the classification task from those that bear no relevant information and cannot derive an appropriate learning rule that allows discriminating among different classes. As a potential solution to this problem, here we exploit the idea of reducing and rotating the feature space in a lower-dimensional gauge and propose the gauge-optimal approximate learning (GOAL) algorithm, which provides an analytically tractable joint solution to the dimension reduction, feature segmentation, and classification problems for small data learning problems. We prove that the optimal solution of the GOAL algorithm consists in piecewise-linear functions in the Euclidean space and that it can be approximated through a monotonically convergent algorithm that presents-under the assumption of a discrete segmentation of the feature space-a closed-form solution for each optimization substep and an overall linear iteration cost scaling. The GOAL algorithm has been compared to other state-of-the-art machine learning tools on both synthetic data and challenging real-world applications from climate science and bioinformatics (i.e., prediction of the El Niño Southern Oscillation and inference of epigenetically induced gene-activity networks from limited experimental data). The experimental results show that the proposed algorithm outperforms the reported best competitors for these problems in both learning performance and computational cost.
Regression learning is one of the long-standing problems in statistics, machine learning, and deep learning (DL). We show that writing this problem as a probabilistic expectation over (unknown) feature probabilities – thus increasing the number of unknown parameters and seemingly making the problem more complex—actually leads to its simplification, and allows incorporating the physical principle of entropy maximization. It helps decompose a very general setting of this learning problem (including discretization, feature selection, and learning multiple piece-wise linear regressions) into an iterative sequence of simple substeps, which are either analytically solvable or cheaply computable through an efficient second-order numerical solver with a sublinear cost scaling. This leads to the computationally cheap and robust non-DL second-order Sparse Probabilistic Approximation for Regression Task Analysis (SPARTAn) algorithm, that can be efficiently applied to problems with millions of feature dimensions on a commodity laptop, when the state-of-the-art learning tools would require supercomputers. SPARTAn is compared to a range of commonly used regression learning tools on synthetic problems and on the prediction of the El Niño Southern Oscillation, the dominant interannual mode of tropical climate variability. The obtained SPARTAn learners provide more predictive, sparse, and physically explainable data descriptions, clearly discerning the important role of ocean temperature variability at the thermocline in the equatorial Pacific. SPARTAn provides an easily interpretable description of the timescales by which these thermocline temperature features evolve and eventually express at the surface, thereby enabling enhanced predictability of the key drivers of the interannual climate.
With the help of high-performance computing, we benchmarked a selection of machine learning classification algorithms on the tasks of whisker stimulus detection, stimulus classification and behavior prediction based on electrophysiological recordings of layer-resolved local field potentials from the barrel cortex of awake mice. Machine learning models capable of accurately analyzing and interpreting the neuronal activity of awake animals during a behavioral experiment are promising for neural prostheses aimed at restoring a certain functionality of the brain for patients suffering from a severe brain injury. The liquid state machine, a highly efficient spiking neural network classifier that was designed for implementation on neuromorphic hardware, achieved the same level of accuracy compared to the other classifiers included in our benchmark study. Based on application scenarios related to the barrel cortex and relevant for neuroprosthetics, we show that the liquid state machine is able to find patterns in the recordings that are not only highly predictive but, more importantly, generalizable to data from individuals not used in the model training process. The generalizability of such models makes it possible to train a model on data obtained from one or more individuals without any brain lesion and transfer this model to a prosthesis required by the patient. Author Summary A neural prosthesis is a computationally driven device that restores the functionality of a damaged brain region for locked-in patients suffering from the aftereffects of a brain injury or severe stroke. As such devices are chronically implanted, they rely on small, low-powered microchips with limited computational resources. Based on recordings describing the neural activity of awake mice, we show that spiking neural networks, which are especially designed for microchips, are able to provide accurate classification models in application scenarios relevant in neuroprosthetics. Furthermore, models were generalizable across mice, corroborating that it will be possible to train a model on recordings from healthy individuals and transfer it to the patient’s prosthesis.
Financial decision-making problems based on relatively few observations and several explanatory variables can be problematic for the common machine learning (ML) tools, since they cannot efficiently discriminate the relevant information. To investigate the challenges of this “small data” regime, we employ several state-of-the-art ML methods for predicting whether three selected stocks from the Swiss Market Index will outperform the market, by using, as classification features, a set of commonly used technical indicators. We show that the recently introduced entropic Scalable Probabilistic Approximation (eSPA) algorithm significantly surpasses its competitors in both prediction accuracy and computational cost. We then discuss the interpretability of the employed ML methods and suggest some statistically derived heuristics to select the most appropriate and parsimonious financial decision-making candidate model.
We propose a pipeline for synthetic generation of personalized Computer Tomography (CT) images, with a radiation exposure evaluation and a lifetime attributable risk (LAR) assessment. We perform a patient-specific performance evaluation for a broad range of denoising algorithms (including the most popular deep learning denoising approaches, wavelets-based methods, methods based on Mumford–Shah denoising, etc.), focusing both on accessing the capability to reduce the patient-specific CT-induced LAR and on computational cost scalability. We introduce a parallel Probabilistic Mumford–Shah denoising model (PMS) and show that it markedly-outperforms the compared common denoising methods in denoising quality and cost scaling. In particular, we show that it allows an approximately 22-fold robust patient-specific LAR reduction for infants and a 10-fold LAR reduction for adults. Using a normal laptop, the proposed algorithm for PMS allows cheap and robust (with a multiscale structural similarity index >90%) denoising of very large 2D videos and 3D images (with over 107 voxels) that are subject to ultra-strong noise (Gaussian and non-Gaussian) for signal-to-noise ratios far below 1.0. The code is provided for open access.
Classification problems in the small data regime (with small data statistic T and relatively large feature space dimension D) impose challenges for the common machine learning (ML) and deep learning (DL) tools. The standard learning methods from these areas tend to show a lack of robustness when applied to data sets with significantly fewer data points than dimensions and quickly reach the overfitting bound, thus leading to poor performance beyond the training set. To tackle this issue, we propose eSPA+, a significant extension of the recently formulated entropy-optimal scalable probabilistic approximation algorithm (eSPA). Specifically, we propose to change the order of the optimization steps and replace the most computationally expensive subproblem of eSPA with its closed-form solution. We prove that with these two enhancements, eSPA+ moves from the polynomial to the linear class of complexity scaling algorithms. On several small data learning benchmarks, we show that the eSPA+ algorithm achieves a many-fold speed-up with respect to eSPA and even better performance results when compared to a wide array of ML and DL tools. In particular, we benchmark eSPA+ against the standard eSPA and the main classes of common learning algorithms in the small data regime: various forms of support vector machines, random forests, and long short-term memory algorithms. In all the considered applications, the common learning methods and eSPA are markedly outperformed by eSPA+, which achieves significantly higher prediction accuracy with an orders-of-magnitude lower computational cost.
We propose a pipeline for a synthetic generation of personalized Computer Tomography (CT) images, with a radiation exposure evaluation and a lifetime attributable risk (LAR) assessment. We perform a patient-specific performance evaluation for a broad range of denoising algorithms (including the most popular Deep Learning denoising approaches, wavelets-based methods, methods based on Mumford-Shah denoising etc.), focusing both on accessing the capability to reduce the patient-specific CT-induced LAR and on computational cost scalability. We introduce a parallel probabilistic Mumford-Shah denoising model (PMS), showing that it markedly-outperforms the compared common denoising methods in denoising quality and cost scaling. In particular, we show that it allows an approximately 22-fold robust patient-specific LAR reduction for infants and a 10-fold LAR reduction for adults. Using a normal laptop the proposed algorithm for PMS allows a cheap and robust (with the Multiscale Structural Similartity index > 90%) denoising of very large 2D videos and 3D images (with over 10 7 voxels) that are subject to ultra-strong Gaussian and various non-Gaussian noises, also for Signal-to-Noise Ratios much below 1.0. The code is provided for open access. One-sentence summary Probabilisitc formulation of Mumford-Shah principle (PMS) allows a cheap quality-preserving denoising of ultra-noisy 3D images and 2D videos.
O. Schenk合作论文数Computer Science Department2