Functional magnetic resonance imaging (fMRI) offers a rich source of data for studying the neural basis of cognition. Here, we describe the Brain Imaging Analysis Kit (BrainIAK), an open-source, free Python package that provides computationally optimized solutions to key problems in advanced fMRI analysis. A variety of techniques are presently included in BrainIAK: intersubject correlation (ISC) and intersubject functional connectivity (ISFC), functional alignment via the shared response model (SRM), full correlation matrix analysis (FCMA), a Bayesian version of representational similarity analysis (BRSA), event segmentation using hidden Markov models, topographic factor analysis (TFA), inverted encoding models (IEMs), an fMRI data simulator that uses noise characteristics from real data (fmrisim), and some emerging methods. These techniques have been optimized to leverage the efficiencies of high-performance compute (HPC) clusters, and the same code can be se amlessly transferred from a laptop to a cluster. For each of the aforementioned techniques, we describe the data analysis problem that the technique is meant to solve and how it solves that problem; we also include an example Jupyter notebook for each technique and an annotated bibliography of papers that have used and/or described that technique. In addition to the sections describing various analysis techniques in BrainIAK, we have included sections describing the future applications of BrainIAK to real-time fMRI, tutorials that we have developed and shared online to facilitate learning the techniques in BrainIAK, computational innovations in BrainIAK, and how to contribute to BrainIAK. We hope that this manuscript helps readers to understand how BrainIAK might be useful in their research.
In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook. Many of our ML workloads have unique characteristics, such as sparse memory accesses, large model sizes, as well as high compute, memory and network bandwidth requirements. We co-designed a high-performance, energy-efficient inference accelerator platform based on these requirements. We describe the inference accelerator platform ecosystem we developed and deployed at Facebook: both hardware, through Open Compute Platform (OCP), and software framework and tooling, through Pytorch/Caffe2/Glow. A characteristic of this ecosystem from the start is its openness to enable a variety of AI accelerators from different vendors. This platform, with six low-power accelerator cards alongside a single-socket host CPU, allows us to serve models of high complexity that cannot be easily or efficiently run on CPUs. We describe various performance optimizations, at both platform and accelerator level, which enables this platform to serve production traffic at Facebook. We also share deployment challenges, lessons learned during performance optimization, as well as provide guidance for future inference hardware co-design.
Deep learning (DL) is one of the most prominent branches of machine learning. Due to the immense computational cost of DL workloads, industry and academia have developed DL libraries with highly-specialized kernels for each workload/architecture, leading to numerous, complex code-bases that strive for performance, yet they are hard to maintain and do not generalize. In this work, we introduce the batch-reduce GEMM kernel and show how the most popular DL algorithms can be formulated with this kernel as the basic building-block. Consequently, the DL library-development degenerates to mere (potentially automatic) tuning of loops around this sole optimized kernel. By exploiting our new kernel we implement Recurrent Neural Networks, Convolution Neural Networks and Multilayer Perceptron training and inference primitives in just 3K lines of high-level code. Our primitives outperform vendor-optimized libraries on multi-node CPU clusters, and we also provide proof-of-concept CNN kernels targeting GPUs. Finally, we demonstrate that the batch-reduce GEMM kernel within a tensor compiler yields high-performance CNN primitives, further amplifying the viability of our approach.
In this document, we describe LDBC Graphalytics, an industrial-grade benchmark for graph analysis platforms. The main goal of Graphalytics is to enable the fair and objective comparison of graph analysis platforms. Due to the diversity of bottlenecks and performance issues such platforms need to address, Graphalytics consists of a set of selected deterministic algorithms for full-graph analysis, standard graph datasets, synthetic dataset generators, and reference output for validation purposes. Its test harness produces deep metrics that quantify multiple kinds of systems scalability, weak and strong, and robustness, such as failures and performance variability. The benchmark also balances comprehensiveness with runtime necessary to obtain the deep metrics. The benchmark comes with open-source software for generating performance data, for validating algorithm results, for monitoring and sharing performance data, and for obtaining the final benchmark result as a standard performance report.
Machine learning methods are used to improve the efficiency by which turbulence-resolved simulations predict whether a hydrogen jet in air crossflow will successfully ignite. The flush-mounted jet issues perpendicularly from the wall of a low-speed wind tunnel into a turbulent boundary layer wherein a laser-induced optical breakdown (LIB) hotspot is deposited. A detailed hydrogen chemical mechanism is used to model the radicals and any subsequent chemical reactions. A dielectric-barrier discharge actuator generates body forces and hydrogen radicals near the jet orifice. We focus on the success or not of the ignition based on LIB location. A challenge is that definitive determination of this requires long simulations, up to 440 mu s after the LIB deposition. This is particularly expensive since multiple simulations are required to find the threshold. To reduce the computational effort, three short-time (91 mu s) criteria are proposed, evaluated, and compared: a constructed criterion based on detailed observations of radicals near the stoichiometric surface and two machine learning approaches, each trained on 38 realizations. The constructed criterion provides a low-cost estimate of the ignition boundary that is unambiguous in 45 of the 50 training and test trials, so only 10% of them would need to be simulated longer. The trained neural networks correctly predict outcomes in all cases evaluated, with the more automated procedure- a convolutional neural network (CNN) trained on two-dimensional images-providing the most definitive outcome prediction. From the CNN, a sensitivity analysis is used to determine which kernel features from the two-dimensional input data, as defined by intermediate-layer network weights, are important for identifying ignition. (C) 2019 The Combustion Institute. Published by Elsevier Inc. All rights reserved.
Implementing the Iterative Soft-Thresholding Algorithm (ISTA) of compressed sensing for MRI image reconstruction is a good candidate for designing accelerators because real-time functional MRI applications require intensive computations. A straightforward mapping of the computation graph of ISTA onto an FPGA, with a wide enough datapath to saturate memory bandwidth, would require substantial resources, such that a modest size FPGA would not fit the reconstruction pipeline for an entire MRI image. This paper proposes several methods to design the kernel components of ISTA, such as matrix transpose, datapath reuse, parallelism within maps, and data buffering to overcome the problem. Our implementation with Intel OpenCL SDK and performance evaluation on Intel HARPv2 show that our methods can map the reconstruction for the entire 256x256 MRI image with 8 or more channels to its FPGA, while achieving good overall performance.
Domain specific accelerators present new challenges for code generation onto novel instruction sets, communication fabrics, and memory architectures. We introduce a shared intermediate representation to describe both deep learning programs and hardware capabilities, then formulate and apply instruction mapping to determine how a computation can be performed on a hardware system. Our scheduler chooses a specific mapping and determines data movement and computation order. With this system, we demonstrate automated extraction of matrix multiplication kernels from recent deep learning operations. We demonstrate 2--5X better performance on GEMM and GRU execution versus state-of-the-art on new hardware and up to 85% of state-of-the-art performance on existing hardware.
BACKGROUND:MRI is commonly used to evaluate pediatric musculoskeletal pathologies, but same-day/near-term scheduling and short exams remain challenges.PURPOSE:To investigate the feasibility of a targeted rapid pediatric knee MRI exam, with the goal of reducing cost and enabling same-day MRI access.STUDY TYPE:A cost effectiveness study done prospectively.SUBJECTS:Forty-seven pediatric patients.FIELD STRENGTH/SEQUENCE:3T. The 10-minute protocol was based on T2 Shuffling, a four-dimensional acquisition and reconstruction of images with variable T2 contrast, and a T1 2D fast spin-echo (FSE) sequence. A distributed, compressed sensing-based reconstruction was implemented on a four-node high-performance compute cluster and integrated into the clinical workflow.ASSESSMENT:In an Institutional Review Board-approved study with informed consent/assent, we implemented a targeted pediatric knee MRI exam for assessing pediatric knee pain. Pediatric patients were subselected for the exam based on insurance plan and clinical indication. Over a 2-year period, 47 subjects were recruited for the study and 49 MRIs were ordered. Date and time information was recorded for MRI referral, registration, and completion. Image quality was assessed from 0 (nondiagnostic) to 5 (outstanding) by two readers, and consensus was subsequently reached.STATISTICAL TESTS:A Wilcoxon rank-sum test assessed the null hypothesis that the targeted exam times compared with conventional knee exam times were unchanged.RESULTS:Of the 49 cases, 20 were completed on the same day as exam referral. Median time from registration to exam completion was 18.7 minutes. Median reconstruction time for T2 Shuffling was reduced from 18.9 minutes to 95 seconds using the distributed implementation. Technical fees charged for the targeted exam were one-third that of the routine clinical knee exam. No subject had to return for additional imaging.DATA CONCLUSION:The targeted knee MRI exam is feasible and reduces the imaging time, cost, and barrier to same-day MRI access for pediatric patients.LEVEL OF EVIDENCE:2 Technical Efficacy: Stage 6 J. Magn. Reson. Imaging 2019.
Domain specific accelerators present new challenges for code generation onto novel instruction sets, communication fabrics, and memory architectures. We introduce a shared intermediate representation to describe both deep learning programs and hardware capabilities, then formulate and apply instruction mapping to determine how a computation can be performed on a hardware system. Our scheduler chooses a specific mapping and determines data movement and computation order.
Ein Prozessor enthalt ein Front-End, um einen Befehl zu empfangen, einen Decodierer, um den Befehl zu decodieren, eine Mengenoperations-Logikeinheit (SOLU), um den Befehl auszufuhren, und eine Stilllegungseinheit. Die SOLU enthalt eine Logik, um eine erste Menge von Schlussel-Wert-Paaren in einer inhaltsassoziativen Datenstruktur zu speichern, um eine zweite Menge von Schlussel-Wert-Paaren zu empfangen und um die Schlussel-Wert-Paare in den beiden Mengen mit zusammenpassenden Schlusseln zu identifizieren. Die SOLU enthalt eine Logik, um die zweite Menge von Schlussel-Wert-Paaren zu der ersten Menge hinzuzufugen, um eine Ausgangsmenge zu erzeugen, und um eine Operation auf die Werte der Schlussel-Wert-Paare mit zusammenpassenden Schlusseln anzuwenden, die einen einzigen Wert fur den zusammenpassenden Schlussel erzeugt. Die SOLU enthalt eine Logik, um eine Ausgangsmenge zu erzeugen, die die Schlussel-Wert-Paare von der ersten Menge mit zusammenpassenden Schlusseln enthalt, und um die Schlussel-Wert-Paare von der ersten Menge mit eindeutigen Schlusseln zu verwerfen.
Using adoptive cell transfer to reintroduce engineered T cells into patients is an exciting new form of cancer immunotherapy. T cells play a central role in adaptive immunity by eliminating virus‐infected and tumor cells. T cell receptors (TCRs) facilitate this by recognizing antigenic peptides presented by major histocompatibility complexes (MHCs), triggering an immune response. While TCRs in gene‐modified T cells can be used for anti‐tumor clinical responses, clinical trials have shown adverse effects attributable to off‐target toxicity, which stems from the inherent cross‐reactivity of a TCR repertoire developed to engage the larger universe of potential peptide antigens. To understand and begin to address this challenge in immunotherapy, elucidation of TCR structural properties and biochemical characteristics is required.TIL 1383I is a T cell receptor (TCR) that targets the human tyrosinase peptide (hTyr) overexpressed in primary and metastatic melanoma. A recent phase I clinical trial of TIL 1383I using TCR‐gene modified T cells to treat melanoma resulted in one case of durable cancer remission but other cases with partial to no responses. To help understand and possibly engineer improved variants of TIL 1383I, we recombinantly expressed TIL 1383I and crystallized it in complex with its binding partner hTyr presented by HLA‐A2. We then used X‐ray crystallography to solve the structure. The structure provides insight into the binding geometry and amino acid residues in the interface between TIL 1383I and hTyr/HLA‐A2. We have also used surface plasmon resonance to determine the TIL 1383I binding affinity. From this work, we aim to design TCR variants with the goal of improving affinity and specificity toward the hTyr antigen.Support or Funding InformationHarper Cancer Research Institute, National Institutes of Health, National Science Foundation, American Cancer SocietyThis abstract is from the Experimental Biology 2018 Meeting. There is no full text article associated with this abstract published in The FASEB Journal.
Magnetic resonance imaging is capable of producing volumetric images without ionizing radiation. Nonetheless, long acquisitions lead to prohibitively long exams. Compressed sensing (CS) can enable faster scanning via sub-sampling with reduced artifacts. However, CS requires significantly higher reconstruction computation, limiting current clinical applications to 2D/3D or limited-resolution dynamic imaging. Here we analyze the practical limitations to T2 Shuffling, a four-dimensional CS-based acquisition, which provides sharp 3D-isotropic-resolution and multi-contrast images in a single scan. Our improvements to the pipeline on a single machine provide a 3x overall reconstruction speedup, which allowed us to add algorithmic changes improving image quality. Using four machines, we achieved additional 2.1x improvement through distributed parallelization. Our solution reduced the reconstruction time in the hospital to 90 seconds on a 4-node cluster, enabling its use clinically. To understand the implications of scaling this application, we simulated running our reconstructions with a multiple scanner setup typical in hospitals.
Querying the content of images, video, and other non-textual data sources requires expensive content extraction methods. Modern extraction techniques are based on deep convolutional neural networks (CNNs) and can classify objects within images with astounding accuracy. Unfortunately, these methods are slow, needing several milliseconds per image using modern GPUs. The cost of content-based queries over a huge video corpus is prohibitive. A promising approach to reduce the runtime cost of queries of visual content is to use a hierarchical model, such as a cascade, where simple cases are handled by an inexpensive classifier. Prior work has sought to design cascades that optimize the computational cost of inference by, for example, using smaller CNNs. However, we observe that there are critical factors besides the inference time that dramatically impact the overall query time. Notably, by treating the input image format and the requisite data handling costs as part of our query optimization, we can enable much more efficient cascades. We find that by jointly optimizing the CNN architecture and input representation, we can provide up to a 40x speedup over the cascades used in the NoScope video query system. We find up to a 156x speedup over ResNet with no accuracy loss and a nearly 300x speedup for users willing to sacrificing some accuracy.
Despite recent advances in dental radiography, the diagnostic accuracies for some of the most common dental diseases have not improved significantly, and in some cases remain low. Intraoral x-ray is the most commonly used x-ray diagnostic tool in dental clinics. It however suffers from the typical limitations of a 2D imaging modality including structure overlap. Cone-beam computed tomography (CBCT) uses high radiation dose and suffers from image artifacts and relatively low resolution. The purpose of this study is to investigate the feasibility of developing a stationary intraoral tomosynthesis (s-IOT) using spatially distributed carbon nanotube (CNT) x-ray array technology, and to evaluate its diagnostic accuracy compared to conventional 2D intraoral x-ray. A bench-top s-IOT device was constructed using a linear CNT based X-ray source array and a digital intraoral detector. Image reconstruction was performed using an iterative reconstruction algorithm. Studies were performed to optimize the imaging configuration. For evaluation of s-IOT's diagnostic accuracy, images of a dental quality assurance phantom, and extracted human tooth specimens were acquired. Results show s-IOT increases the diagnostic sensitivity for caries compared to intraoral x-ray at a comparable dose level.
Apache Spark is a popular framework for data analytics with attractive features such as fault tolerance and interoperability with the Hadoop ecosystem. Unfortunately, many analytics operations in Spark are an order of magnitude or more slower compared to native implementations written with high performance computing tools such as MPI. There is a need to bridge the performance gap while retaining the benefits of the Spark ecosystem such as availability, productivity, and fault tolerance. In this paper, we propose a system for integrating MPI with Spark and analyze the costs and benefits of doing so for four distributed graph and machine learning applications. We show that offloading computation to an MPI environment from within Spark provides 3.1−17.7× speedups on the four sparse applications, including all of the overheads. This opens up an avenue to reuse existing MPI libraries in Spark with little effort.
The duality between graphs and matrices means that many common graph analyses can be expressed with primitives such as generalized sparse matrix-vector multiplication (SpMSpV) and sparse matrix-matrix multiplication (SpGEMM). Achieving high performance on these primitives is challenging due to limited arithmetic intensity, irregular memory accesses, and significant network communication requirements in the distributed setting. In this paper we implement four graph applications using GraphPad, our optimized multinode implementations of generalized linear algebra primitives such as SpMSpV and SpGEMM. GraphPad is highly flexible to accommodate multiple data layouts, partitioning strategies, and incorporates communication optimizations. Our performance at scale can exceed that of CombBLAS by up to 40×. In addition to GraphPad's performance in a distributed setting, it is also within 2× the performance of GraphMat, a high performance graph framework on a single node for four out of five benchmarks. We also show our communication optimizations and flexibility are critical for good performance on both HPC clusters and commodity cloud platforms.
The scale of functional magnetic resonance image data is rapidly increasing as large multi-subject datasets are becoming widely available and high-resolution scanners are adopted. The inherent low-dimensionality of the information in this data has led neuroscientists to consider factor analysis methods to extract and analyze the underlying brain activity. In this work, we consider two recent multi-subject factor analysis methods: the Shared Response Model and the Hierarchical Topographic Factor Analysis. We perform analytical, algorithmic, and code optimization to enable multi-node parallel implementations to scale. Single-node improvements result in 99χ and 2062x speedups on the two methods, and enables the processing of larger datasets. Our distributed implementations show strong scaling of 3.3x and 5.5χ respectively with 20 nodes on real datasets. We demonstrate weak scaling on a synthetic dataset with 1024 subjects, equivalent in size to the biggest fMRI dataset collected until now, on up to 1024 nodes and 32,768 cores.
T-cell receptors (TCRs) have emerged as a new class of therapeutics, most prominently for cancer where they are the key components of new cellular therapies as well as soluble biologics. Many studies have generated high affinity TCRs in order to enhance sensitivity. Recent outcomes, however, have suggested that fine manipulation of TCR binding, with an emphasis on specificity may be more valuable than large affinity increments. Structure-guided design is ideally suited for this role, and here we studied the generality of structure-guided design as applied to TCRs. We found that a previous approach, which successfully optimized the binding of a therapeutic TCR, had poor accuracy when applied to a broader set of TCR interfaces. We thus sought to develop a more general purpose TCR design framework. After assembling a large dataset of experimental data spanning multiple interfaces, we trained a new scoring function that accounted for unique features of each interface. Together with other improvements, such as explicit inclusion of molecular flexibility, this permitted the design new affinity-enhancing mutations in multiple TCRs, including those not used in training. Our approach also captured the impacts of mutations and substitutions in the peptide/MHC ligand, and recapitulated recent findings regarding TCR specificity, indicating utility in more general mutational scanning of TCR-pMHC interfaces.
We propose BlackOut, an approximation algorithm to efficiently train massive recurrent neural network language models (RNNLMs) with million word vocabularies. BlackOut is motivated by using a discriminative loss, and we describe a new sampling strategy which significantly reduces computation while improving stability, sample efficiency, and rate of convergence. One way to understand BlackOut is to view it as an extension of the DropOut strategy to the output layer, wherein we use a discriminative training loss and a weighted sampling scheme. We also establish close connections between BlackOut, importance sampling, and noise contrastive estimation (NCE). Our experiments, on the recently released one billion word language modeling benchmark, demonstrate scalability and accuracy of BlackOut; we outperform the state-of-the art, and achieve the lowest perplexity scores on this dataset. Moreover, unlike other established methods which typically require GPUs or CPU clusters, we show that a carefully implemented version of BlackOut requires only 1-10 days on a single machine to train a RNNLM with a million word vocabulary and billions of parameters on one billion words. Although we describe BlackOut in the context of RNNLM training, it can be used to any networks with large softmax output layers.