Sequence alignment is an important tool in bioinformatics and computational biology. It uses dynamic programming (DP)-based algorithm to obtain optimal scores during the sequence homology search. This algorithm guarantees for accurate search, however with expense of quadratic time complexity. Thus, researchers have implemented the DP algorithm in Field Programmable Gate Array (FPGA)-based platform. However, the configuration stage also endures several challenges especially for protein sequence alignment. Prior to the sequence homology search, the processing element (PE) requires frequent memory load and rapid access to substitution matrix coefficients. The efficient supply of configuration data for the PE is crucial as to reduce the configuration time, hence affected speed performance of the core system. Typical PE configuration scheme uses serial configuration chain where it configures different look-up tables in the pipeline of PEs sequentially. Consequently, the configuration time increases proportionally to the number PEs. Thus, in this work, a new architecture of PE parallel loader with parallel configuration chain technique has been proposed. The parallel loader consists of several circular buffers, designed using n-bit registers and transmitted to the PEs via large data bus. This allows efficient and simultaneous supply of the configuration data to all PEs. This loader has been implemented on Virtex-5 FPGA and achieved 480.25 MHz clock frequency. It utilized only 52 or 0.3 percent of the XC5VLX110 Virtex-5 slices. Moreover, the parallel loader element length is parameterizable, thus it can load any size of substitution matrix score either BLOSUM or PAM series.
The support vector machine (SVM) is one of the highly powerful classifiers that have been shown to be capable of dealing with high-dimensional data. However, its complexity increases requirements of computational power. Recent technologies including the postgenome data of high-dimensional nature add further complexity to the construction of SVM classifiers. In order to overcome this problem, hardware implementations of the SVM classifier have been proposed to benefit from parallelism to accelerate the SVM. On the other hand, those implementations offer limited flexibility in terms of changing parameters and require the reconfiguration of the whole device. The latter interrupts the operation of other tasks placed on the hardware device. In this work, two flexible hardware implementations of the SVM classifier are proposed, namely A1 and A2 classifiers with successful applications in a microarray dataset. In addition, two dynamically and partially reconfigurable (DPR) architectures of the SVM classifier are presented. The A1 and A2 architectures have achieved up to 61x and 49x speed-up, respectively, over the equivalent general purpose processor. Furthermore, the DPR implementations achieved at least similar to 8x reduction in reconfiguration time compared to non-DPR implementation. This is a significant achievement that can be easily adapted in other application domains of a similar nature.
This article presents a new solution for easing the development of reconfigurable applications using Field-Programable Gate Arrays (FPGAs). Namely, our Reliable Reconfigurable Real-Time Operating System (R3TOS) provides OS-like support for partially reconfigurable FPGAs. Unlike related works, R3TOS is founded on the basis of resource reusability and computation ephemerality. It makes intensive use of reconfiguration at very fine FPGA granularity, keeping the logic resources used only while performing computation and releasing them as soon as it is completed. To achieve this goal, R3TOS goes beyond the traditional approach of using reconfigurable slots with fixed boundaries interconnected by means of a static communication infrastructure. Instead, R3TOS approaches a static route-free system where nearly everything is reconfigurable. The tasks are concatenated to form a computation chain through which partial results naturally flow, and data are exchanged among remotely located tasks using FPGA’s reconfiguration mechanism or by means of “removable” routing circuits. In this article, we describe the R3TOS microkernel architecture as well as its hardware abstraction services and programming interface. Notably, the article presents a set of novel circuits and mechanisms to overcome the limitations and exploit the opportunities of Xilinx reconfigurable technology in the scope of hardware multitasking and dependability.
Bioinformatics data tend to be highly dimensional in nature thus impose significant computational demands. To resolve limitations of conventional computing methods, several alternative high performance computing solutions have been proposed by scientists such as Graphical Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). The latter have shown to be efficient and high in performance. In recent years, FPGAs have been benefiting from dynamic partial reconfiguration (DPR) feature for adding flexibility to alter specific regions within the chip. This work proposes combing the use of FPGAs and DPR to build a dynamic multi-classifier architecture that can be used in processing bioinformatics data. In bioinformatics, applying different classification algorithms to the same dataset is desirable in order to obtain comparable, more reliable and consensus decision, but it can consume long time when performed on conventional PC. The DPR implementation of two common classifiers, namely support vector machines (SVMs) and K-nearest neighbor (KNN) are combined together to form a multi-classifier FPGA architecture which can utilize specific region of the FPGA to work as either SVM or KNN classifier. This multi-classifier DPR implementation achieved at least ~8x reduction in reconfiguration time over the single non-DPR classifier implementation, and occupied less space and hardware resources than having both classifiers. The proposed architecture can be extended to work as an ensemble classifier.
This article describes the contributions of the reliable reconfigurable real-time operating system (r3tos) for building an autonomous fault-tolerant system using currently available xilinx partially reconfigurable field-programmable gate arrays. The authors present an r3tos-based inverter controller of a real-world railway traction system that is proven to recover from most of the errors provoked to it without requiring any human intervention.
One of the most challenging tasks in sequence alignment is its repetitive and time-consuming alignment matrix computations. In addition, performing sequence alignment in hardware, i.e. FPGA requires more hardware resources as the number of processing elements is replicated to increase performance throughput. This paper first reviews the existing FPGA-based biological sequence alignment core architectures and then proposed an efficient scheduling strategy, the so-called overlap computation and configuration (OCC) towards realizing optimized biological sequence alignment core architecture targeting for pairwise sequence alignment. In this research work, double buffering-based core architecture have been proposed and implemented on Xilinx Virtex-5 FPGA. Results have shown that this approach gained more than 10K times speed-up as compared to the GPP solution.
Despite the clear potential of FPGAs to push the current power wall beyond what is possible with general-purpose processors, as well as to meet ever more exigent reliability requirements, the lack of standard tools and interfaces to develop reconfigurable applications limits FPGAs' user base and makes their programming not productive. R3TOS is our contribution to tackle this problem. It provides systematic OS support for FPGAs, allowing the exploitation of some of the most advanced capabilities of FPGA technology by inexperienced users. What makes R3TOS special is its nonconventional way of exploiting on-chip resources: These are used indistinguishably for carrying out either computation or communication tasks at different times. Indeed, R3TOS does not rely on any static infrastructure apart from its own core circuitry, which is constrained to a specific region within the FPGA where it is implemented. Thus, the rest of the device is kept free of obstacles, with the spare resources ready to be used as and whenever needed. At runtime, the hardware tasks are scheduled and allocated with the dual objective of improving computation density and circumventing damaged resources on the FPGA.
This special section of IEEE Transactions on Computers presents some of the latest research developments in the field of adaptive hardware and systems. The creation of this section was motivated by lively discussions held at the annual NASA/ESA Adaptive Hardware and Systems (AHS) conference, which showed a need for such special section at a top ranked journal. At the end of a rigorous review process, ten papers were selected for publication from a set of high quality submissions consisting of regular papers and extended papers from the AHS 2012 conference proceedings. The articles are then briefly described.
This paper presents a novel miniaturized reconfigurable and switchable feeding network to cover GSM, GPS, 3G, WiFi and global LET standards. The feeding network consists of four conventional Wilkinson power dividers which can be individually reconfigured in length using PIN diodes switches. By controlling the bias voltages of these PIN diodes, the operating frequency of the proposed design can be converted between four different bands: 600MHz-900MHz, 1.2GHz-1.6GHz, 1.8GHz-2.2GHz and 2.4GHz-2.6GHz. The first frequency band (600MHz-900MHz) is applied to satisfy the applications of LTE US (700MHz), LTE UK (800MHz) and GSM (850MHz, 900MHz). The second band (1.2GHz-1.6GHz) targets GPS L1 (1.575GHz) and GPS L2 (1.227GHz). Different GSM (1800MHz, 1900MHz) and 3G standards (UMTS, W-CDMA, TD-SCDMA and CDMA2000) are located in the third frequency band (1.8GHz-2.2GHz). The last band (2.4GHz-2.6GHz) is used to cover WiFi (2.45GHz) and LTE Europe (2.6GHz). The miniaturized and optimized feeding network exhibits good performance for S-Parameters in each band, which includes low return loss, equal power splitting and suitable insertion loss. Within the simulation environment, three types of PIN diode models were constructed and investigated in order to improve accuracy. The feeding network is implemented on an FR4 substrate. Fabrication and measurement results closely correlate with those obtained during design simulations. The reconfigurable feeding network can be particularly applied to commercial multiband communication systems.
This paper addresses the high-performance systems which are based on swapping relocatable partial bitstreams (also called hardware tasks) in and out of an FPGA device using Dynamic Partial Reconfiguration (DPR) in Xilinx Virtex FPGAs. Configuration speed is important in such systems to achieve high performance. Previous research efforts were focused on over-clocking the Internal Configuration Access Port (ICAP) and compressing the partial bitstreams of the hardware tasks to enhance the configuration speed. We propose the use of the Multiple Frame Write (MFW) feature to significantly reduce the configuration time when multiple instances of a hardware task are needed. In this paper, we demonstrate the design and implementation of a novel internal reconfiguration engine which dynamically generates partial bitstreams required for simultaneous configuration of multiple clones of relocatable hardware tasks.
The current trend in embedded vision systems is to propose bespoke solutions for specific problems as each application has different requirement and constraints. There is no widely used model or benchmark which aims to facilitate generic solutions in embedded vision systems. Providing such model is a challenging task due to the wide number of use cases, environmental factors, and available technologies. However, common characteristics can be identified to propose an abstract model. Indeed, the majority of vision applications focus on the detection, analysis and recognition of objects. These tasks can be reduced to vision functions which can be used to characterize the vision systems. In this paper, we present the results of a thorough analysis of a large number of different types of vision systems. This analysis led us to the development of a system’s taxonomy, in which a number of vision functions as well as their combination characterize embedded vision systems. To illustrate the use of this taxonomy, we have tested it against a real vision system that detects magnetic particles in a flowing liquid to predict and avoid critical machinery failure. The proposed taxonomy is evaluated by using a quantitative parameter which shows that it covers 95 percent of the investigated vision systems and its flow is ordered for 60 percent systems. This taxonomy will serve as a tool for classification and comparison of systems and will enable the researchers to propose generic and efficient solutions for same class of systems.
Advancements in silicon, software and IP support have made Field Programmable Gate Arrays (FPGAs) a highly flexible solution for many applications. With the growing number of companies providing IP support for FPGAs, IP license violations by over-deployment of IP into more devices than originally licensed remains a major concern for IP owners. In this paper we present a solution for secure IP exchange and configuration based on the Dynamic Partial Reconfiguration (DPR) feature in Xilinx FPGAs. Our system deploys DPR to integrate encrypted hard-macro IP cores into identifiable FPGA devices. These IP cores are configured using a proposed partial bitstream relocation technique to allow for a flexible design flow. We present a proof-of-concept implementation of a secure internal reconfiguration engine on a Xilinx Virtex-6 FPGA.
Filters datapath signals and coefficients are quantised when implemented in hardware to limit and reduce excessive hardware requirements. In this paper, we formulate quantisation errors (noises) in fixed-point arithmetic using a novel analytical model. The latter extends a conventional signal quantisation statistical model by assuming a Gaussian distribution noise. The paper gives the mathematical expressions to compute the statistical parameters and range values of the quantisation errors at any point in a multistage FIR filters structure depending on the wordlengths fractional precisions. Three case studies are included to vindicate the model@?s validity and accuracy in predicting the quantisation error parameters in the absence of filter coefficients quantisation. To counter the effects of the latter, we present a novel approach, called errors cancellation. The approach tends to represent the filter coefficients using different wordlengths to minimise the dynamic of the error filter@?s output. This allows limiting the quantisation effects to the signals quantisation only, which is statistically accurately modelled. The validity and the efficiency of the approach along with our analytical model are shown using two further case studies. Through our errors cancellation approach and analytical model, a hardware designer can now minimise the effects of the filter coefficients quantisation and predict subsequently the range values of the computation errors depending on the fractional precision used. He can also preset the latter to achieve the sought computations accuracy.
This paper describes a novel way to exploit the computation capabilities delivered by modern Field-Programmable Gate Arrays (FPGAs), not only towards a higher performance, but also towards an improved reliability. Computation-specific pieces of circuitry are dynamically scheduled and allocated to different resources on the chip based on a set of novel algorithms which are described in detail in this article. These algorithms consider most of the technological constraints existing in modern partially reconfigurable FPGAs as well as spontaneously occurring faults and emerging permanent damage in the silicon substrate of the chip. In addition, the algorithms target other important aspects such as communications and synchronization among the different computations that are carried out, either concurrently or at different times. The effectiveness of the proposed algorithms is tested by means of a wide range of synthetic simulations, and, notably, a proof-of-concept implementation of them using real FPGA hardware is outlined.
Classifying Microarray data, which are of high dimensional nature, requires high computational power. Support Vector Machines-based classifier (SVM) is among the most common and successful classifiers used in the analysis of Microarray data but also requires high computational power due to its complex mathematical architecture. Implementing SVM on hardware exploits the parallelism available within the algorithm kernels to accelerate the classification of Microarray data. In this work, a flexible, dynamically and partially reconfigurable implementation of the SVM classifier on Field Programmable Gate Array (FPGA) is presented. The SVM architecture achieved up to 85× speed-up over equivalent general purpose processor (GPP) showing the capability of FPGAs in enhancing the performance of SVM-based analysis of Microarray data as well as future bioinformatics applications.
High-Performance Computing using FPGA covers the area of high performance reconfigurable computing (HPRC). This book provides an overview of architectures, tools and applications for High-Performance Reconfigurable Computing (HPRC). FPGAs offer very high I/O bandwidth and fine-grained, custom and flexible parallelism and with the ever-increasing computational needs coupled with the frequency/power wall, the increasing maturity and capabilities of FPGAs, and the advent of multicore processors which has caused the acceptance of parallel computational models. The Part on architectures will introduce different FPGA-based HPC platforms: attached co-processor HPRC architectures such as the CHRECs Novo-G and EPCCs Maxwell systems; tightly coupled HRPC architectures, e.g. the Convey hybrid-core computer; reconfigurably networked HPRC architectures, e.g. the QPACE system, and standalone HPRC architectures such as EPFLs CONFETTI system. The Part on Tools will focus on high-level programming approaches for HPRC, with chapters on C-to-Gate tools (such as Impulse-C, AutoESL, Handel-C, MORA-C++); Graphical tools (MATLAB-Simulink, NI LabVIEW); Domain-specific languages, languages for heterogeneous computing(for example OpenCL, Microsofts Kiwi and Alchemy projects). The part on Applications will present case from several application domains where HPRC has been used successfully, such as Bioinformatics and Computational Biology; Financial Computing; Stencil computations; Information retrieval; Lattice QCD; Astrophysics simulations; Weather and climate modeling.
Field programmable hardware gives electronic systems the ability to be reconfigured at run time. This allows electronic systems to be more efficiently customized on demand and on-the-fly depending on user requirements and environmental changes. This paper presents a run-time reconfigurable system that allows computing tasks to adjust their sizes in response to current available resources, optimizing the overall performance by maximally exploiting all the resources on the chip. In particular, we present a novel run-time task assembler, which assembles tasks with desired parameters on-the-fly, together with an efficacious run-time task placer to rapidly configure tasks at optimum locations. The system is demonstrated with a dynamic programming-based pairwise sequence alignment application. Real hardware implementation result shows that our run-time reconfigurable system optimizes resource usage on the fly by ~ 3x, while matching the performance of carefully hand-crafted static solutions.