矩阵运算是高性能计算中核心问题之一,矩阵分解是提高矩阵运算并行性的重要途径,飞速发展的FPGA为并行运算结构提供了有力的环境支持.该文基于子矩阵更新同一化算法实现了Cholesky分解,基于FPGA设计了相应的并行结构.实验结果表明:与通用处理器的软件实现相比,本文实现的Cholesky分解的FPGA并行结果在核心计算性能上可以取得10倍以上的加速比,该算法针对矩阵三角化计算过程具有更高的数据和流水并行性.
A triple exposure, high dynamic range (HDR), CMOS image sensor with an active array size of 1280 x 1080, and sub 1enoise floor is presented. This sensor is the first sensor that utilizes OmniVision’s split-pixel technology ported to the BSI process technology. The pixel size is 4.2um, and the pixel incorporates programmable conversion gain (CG). There are three exposure values: the long exposure channel (L), captured by the large photo diode (LPD); the short exposure (S), captured by the small photo diode (SPD); and the very short exposure (VS), captured by the LPD. The three exposure values are A/D converted and processed digitally to generate HDR pixel values with a 20-bits linear range. The sensor is able to output full resolution at 60fps, with both serial MIPI and parallel DVP output being supported.
The reconfigurable computing system became an important choice according to accelerating compute-intensive ap-plications.Among most compute-intensive applications,the matrix triangularization decomposition always was in the central position of research subjects and presented a great value to solve linear equation systems and matrix eigenvalue problems in science or engi-neering area.This paper analyzed the linear computing process of triangularization and proposed a hardware-adaptive parallel sub-matrix identity updating algorithm and a high-performance parallel structure hardware template for matrix triangularization on FPGA (Field Programmable Gate Array)according to the common triangularization computing process of the matrix triangularization de-composition.The research focused on the high-performance FPGA parallel structure implementation and optimization methods for the LU matrix triangularization decomposition.In theoretical analysis,the proposed algorithm presents better pipeline-parallelism and da-ta-parallelism during the matrix triangularization process.The experimental result shows that the proposed structure gets over decuple speedup compared to general-purpose processors and the previous works in vital performance.
A novel rectangular shape differential CMOS split-drain Hall Effect magnetic field-effect transistor (MAGFET) was designed and fabricated employing a CMOS 0.5 μm process. The detection and monitoring of single 2.8 μm diameter magnetic beads was successfully performed using this MAGFET design. Based on the device modeling, simulation and the signal to noise ratio analysis, it was found that the optimal sensitivity can be achieved when the MAGFET channel width to length ratio is equal to 1.3. Further, it is shown through that when the MAGFET is scaled down, its SNR performance can sustain its peak, while being more sensitive to the geometry variations.
A flexible technology is proposed to integrate smart electronics and microfluidics all embedded in an elastomer package. The microfluidic channels are used to deliver both liquid samples and liquid metals to the integrated circuits (ICs). The liquid metals are used to realize electrical interconnects to the IC chip. This avoids the traditional IC packaging challenges, such as wire-bonding and flip-chip bonding, which are not compatible with current microfluidic technologies. As a demonstration we integrated a CMOS magnetic sensor chip and associate microfluidic channels on a polydimethylsiloxane (PDMS) substrate that allows precise delivery of small liquid samples to the sensor. Furthermore, the packaged system is fully functional under bending curvature radius of one centimetre and uniaxial strain of 15%. The flexible integration of solid-state ICs with microfluidics enables compact flexible electronic and lab-on-a-chip systems, which hold great potential for wearable health monitoring, point-of-care diagnostics and environmental sensing among many other applications.
To aid in the hardware/software partitioning of the reconfigurable computing systems, it is necessary to conduct fast and accurate FPGA-based delay estimations before the partitioning. Most previous works predict the delay by adopting a high-level delay estimation based on empirical formulae. In such method, the empirical formulae are often obtained by a regression analysis on the real values reported by the synthesis and place-and-route tools of FPGAs. With alternative properties of tools or different FPGA devices, the empirical formulae need to be re-analyzed and decided. However, it is time-consuming due to inevitably repeated running synthesis and place-and-route tasks, which results in slow estimation and always beyond the tolerance of the estimation time. To address this problem, we present an improved high-level delay-estimation method in this article. We derived theory formulae called increasing formulae for HLL (High Level Language) operations from the basic idea of the hardware circuit design. These increasing formulae can be fit for most FPGAs. Combining the proposed formulae, the paper proposes a rapid estimation algorithm also. And the algorithm can obtain hardware delay of different hardware versions, thus reduces the number of times of running the time-consuming tasks greatly. Experimental results show that our method can achieve error within 2.69% for virtex-5 FPGA, compared with the real values.
To aid in the hardware/software partitioning of recon gurable computing systems, fast yet accurate FPGA based delay estimations are necessary before the partitioning. Most previous works predict the delay by using a high-level delay-estimation method of the empirical formulae. However, this method needs to run many times of the time-consuming synthesis, place and route procedures, which may take up to hours or days for all possible partition options. To address this problem, this paper proposed an auto estimation model to improve the previous high-level delay-estimation. In this model, we rstly derive calculation formulae called increasing formulae of HLL operations from the basic idea of the hardware circuit design. Then the feedback based framework is applied to adjust the increasing formulae for alternative FPGAs or synthesis properties, and estimate the delay of the partitioning. This model reduces the times of running the time-consuming procedures. Experimental results show the method can achieves error within 5% for virtex-5 FPGA, compared with the real delay. © 2013 Binary Information Press.
We have demonstrated flexible packaging and integration of CMOS IC chips with PDMS microfluidics. Microfluidic channels are used to deliver both liquid samples and liquid metals to the CMOS die. The liquid metals are used to realize electrical interconnects to the CMOS chip. As a demonstration we integrated a CMOS magnetic sensor die and matched PDMS microfluidic channels in a flexible package. The packaged system is fully functional under 3cm bending radius. The flexible integration of CMOS ICs with microfluidics enables previously unavailable flexible CMOS electronic systems with fluidic manipulation capabilities, which hold great potential for wearable health monitoring, point-of-care diagnostics and environmental sensing.
We proposed a portable molecular diagnostic tool implemented by CMOS and Microfluidic technology for the early HIV diagnosis. Our system diagnoses HIV by detecting the P24 antigen concentration in blood sample. We labeled P24 antigens with chemiluminescence signal, which can be detected by our CMOS photon detector. We designed the single photon avalanche diode (SPAD) with single molecule sensitivity, which can help to diagnosis HIV patient earlier, thus increases the chance of cure and reduce the change of further infection spread. To make the system reusable, we use the 10 um diameter magnetic beads as the assay substrate, which can be coated with the anti-P24 antibodies. These magnetic beads can be delivered through microfluidic channels to the sensing area. Then they can catch p24 antigens from the blood sample and generate chemiluminescence signal. After the test, the magnetic beads can be easily washed away, and the system can be reused.
Maternal effect genes encode proteins that are produced during oogenesis and play an essential role during early embryogenesis. Genetic ablation of such genes in oocytes can result in female subfertility or infertility. Here we report a newly identified maternal effect gene, Nlrp2, which plays a role in early embryogenesis in the mouse. Nlrp2 mRNAs and their proteins (∼118 KDa) are expressed in oocytes and granulosa cells during folliculogenesis. The transcripts show a striking decline in early preimplantation embryos before zygotic genome activation, but the proteins remain present through to the blastocyst stage. Immunogold electron microscopy revealed that the NLRP2 protein is located in the cytoplasm, nucleus and close to nuclear pores in the oocytes, as well as in the surrounding granulosa cells. Using RNA interference, we knocked down Nlrp2 transcription specifically in mouse germinal vesicle oocytes. The knockdown oocytes could progress through the metaphase of meiosis I and emit the first polar body. However, the development of parthenogenetic embryos derived from Nlrp2 knockdown oocytes mainly blocked at the 2-cell stage. The maternal depletion of Nlrp2 in zygotes led to early embryonic arrest. In addition, overexpression of Nlrp2 in zygotes appears to lead to normal development, but increases blastomere apoptosis in blastocysts. These results provide the first evidence that Nlrp2 is a member of the mammalian maternal effect genes and required for early embryonic development in the mouse.
This paper presents an efficient parallel architecture for fast solving linear system of equations over binary operations of GF(2),which is derived from a proposed hardware-optimized Gaussian elimination.The optimization of the Gaussian elimination with pivot element is realized by using parallel elimination and cyclic shift operations instead of loop nest in each iteration.A mesh structure of "smart memory" cells is proposed for building the whole parallel architecture where the modified algorithm is mapped onto.The average running time of the architecture for n-dimension binary matrix equals 2n cycles as opposed to about 1/4n3 in software.Experimental results show that the performance of the system is improved by about two orders of magnitude.
The loop structure is always considered as the main time-consuming part in most computationally intensive applications.Since the FPGA-based reconfigurable computing systems emerge in recent years,the static techniques for analyzing loop structures are not able to meet the requirement of specific optimization according to the current behavior of programs.To address the lack of directly accessing the run-time information by using the dynamic techniques for analyzing loops,a new loop-analysis method is proposed.In this method which is implemented on the Low Level Virtual Machine(LLVM),the loop structures obtained from the Control Flow Graph(CFG)are recognized according to the dominating relationship,then the result of the edge profiling before the frequency of loop-calling,the average frequency of iteration and time of running are calculated.Experimental results manifest that the proposed method can recognize all the loop structure and collect the loop run-time information accurately,which can support hardware/software partitioning work of reconfigurable computing.
Time-Correlated Single Photon Counting (TCSPC) can provide not only the time information of a photon, but also the photon density information. Based on the conclusion of usual time interval measuring methods, this paper chooses the scheme of Time-to-Digital Converter (TDC) based on delay line structure, meeting the TCSPC system's requirement for high timing resolution. This TDC contains two delay lines and a main counter. After finishing the framework of the TDC using Verilog, we confirm the architecture of the delay element by simulation and on-board test. Using the histogram from FPGA, the TDC system implement is for time resolution below 200 ps.
A novel circular CMOS MAGFET (Magnetic Field Effect Transistor) design is introduced and a novel device geometry design methodology is proposed to optimize the magnetic particle detection sensitivity of such devices. In order to optimize the signal to noise ratio, it was determined that the geometry of the MAGFET is required to have specific ratios, where its sector angle θ and its inner and outer radii r 1 and r 2 are optimized when θ/ ln ( r 2 / r 1 ) = 1.3 . Compared to the more traditional rectangular MAGFET, the circular MAGFET has compatible SNR peak performance with rectangular MAGET. However, when the size of the MAGFET is scaled down in order to detect smaller magnetic particles, the proposed circular MAGFET has more robust SNR performance, design flexibility and tolerance to processing variations.
In this paper, we proposed a parallel hardware methodology employing the modified Gaussian elimination algorithm to efficiently solve linear system of equations (LSEs). Two parallel operators are issued in the hardware-optimized algorithm. Moreover, to be the proof-of-concept, the proposed parallel methodology is implemented to hardware structures in cases to address solving LSEs over GF(2) (primarily are bits operation) and LSEs with floating-point (IEEE-754 standard, 32-bit single precision) coefficient matrix. The corresponding hardware is mainly composed of uniformly distributed basic cells which store and register data, yielding a standalone worst case time complexity O(n 2) opposed to O(n 3) of the software replication. Finally, the given experimental result inosculated with the theory analysis.
A Single-Photon Avalanche Diode (SPAD) design solution that can be implemented in a low cost CMOS process, n-well 0.5μm process, is proposed. This SPAD design used the lateral diffusion of the n-wells to create a low n doping density area as the guard ring to prevent the premature breakdown at the edge of SPAD. Through the TCAD simulation and fabricated design characterization, we found the proper gap length between n-wells to create the guard ring that allows the SPAD to work in the Geiger Mode. The SPAD has dark count rate (DCR) of 750/Hz at 14.85V bias voltage without cooling and 60ns dead time.