Variability analysis is important in successfully deploying multi-gigabit backplane printed wiring boards (PWBs) with growing numbers of high-speed SerDes links. We discuss the need for large sample sizes to obtain accurate variability estimates of SI metrics (eye height, phase skew, etc). Using a dataset of 11,961 S-parameters, we demonstrate statistical techniques to extract accurate estimates of PWB SI performance variations. We cite numerical examples illustrating how these variations may contribute to underestimated or overestimated design criteria, causing unnecessary design expense. Tabular summaries of performance variation and key findings of broad interest to the general SI community are highlighted.
Applications such as public-key cryptography are critically reliant on the speed of modular multiplication for their performance. This paper introduces a new block-based variant of Montgomery multiplication, the Block Product Scanning (BPS) method, which is particularly efficient using new 512-bit advanced vector instructions (AVX-512) on modern Intel processor families. Our parallel-multiplication approach also allows for squaring and sub-quadratic Karatsuba enhancements. We demonstrate $$1.9\,\times $$ improvement in decryption throughput in comparison with OpenSSL and $$1.5\,\times $$ improvement in modular exponentiation throughput compared to GMP-6.1.2 on an Intel Xeon CPU. In addition, we show $$1.4\,\times $$ improvement in decryption throughput in comparison with state-of-the-art vector implementations on many-core Knights Landing Xeon Phi hardware. Finally, we show how interleaving Chinese remainder theorem-based RSA calculations within our parallel BPS technique halves decryption latency while providing protection against fault-injection attacks.
Optoacoustic, or photoacoustic, imaging combines the penetration capabilities of ultrasound imaging with the contrast mechanism of optical absorption related to the photoacoustic effect. To enable modeling of photoacoustic measurements and imaging applications, the problem can be divided into modeling of the optical photon propagation and the resulting acoustic wave propagation. In highly scattering media such as soft tissues, Monte Carlo (MC) methods are used to evaluate photon propagation, absorption and the subsequent generation of a heat source that produces the photoacoustic effect. Acoustic wave equations are then used to model the propagation of the sound in the medium. However, if the appropriate physics could be modeled with a single tool, there would be advantages of consistency in spatial and temporal domains as well as easier integration of the optical and acoustic simulations. We propose using MC methods for optical photon propagation and for the acoustic wave propagation. The simulations are split into two stages for the optical photon propagation and the acoustic phonon propagation. We describe how the same MC framework is used to model photon propagation and acoustic phonon propagation. We will then demonstrate this new combined MC framework in models of photoacoustic problems in homogeneous and heterogeneous cases. These results will be compared with results from using k-Wave models for the acoustic wave propagation. The correspondence between k-Wave the acoustic MC model are shown to be in good agreement.
As network data rates advance toward 1 Tb/s, hardware-based implementations of anti-replay offer desirable tradeoffs over software. However, internal logic busses in FPGAs are becoming wider (512+ bits) and segmented (more than one packet per clock cycle) to accommodate increased network data rates. Such busses are challenging for applications such as anti-replay that require read-modify-write operations to a coherent database on each packet arrival. In this paper we present an FPGA-targeted pipelined anti-replay design capable of accommodating 1024 IPsec tunnels at 1 Tb/s data rate. The novel design is enabled by fast on-chip block RAMs in a xcvu190 Virtex Ultrascale FPGA that are used to construct a 20-port RAM memory operating at 400 MHz with over 5 Tb/s of peak bandwidth. Custom single-clock write-combining techniques are described that accommodate multiple concurrent updates to the same database address. We also investigate the limits of capacity and concurrency for the anti-replay application.
The Advanced Encryption Standard (AES) together with the Galois Counter Mode (GCM) of operation has been approved for use in several high throughput network protocols to provide authenticated encryption. However, the demand for continued increase in network bandwidth has not abated and we anticipate the need for continual performance improvement of AES-GCM in hardware. Additionally, as data interfaces become wider and segmented, existing methods of GCM parallelization become inefficient. This paper presents a novel scalable architecture for highly parallel implementations of AES-GCM that can process multiple separately-keyed packets simultaneously every clock cycle. We demonstrate throughputs of 482 Gb/s in a single Xilinx Virtex Ultrascale FPGA and describe how the architecture can be used to achieve over 800 Gb/s in a system comprising multiple FPGAs.
Embedded microcontroller applications often experience multiple limiting constraints: memory, speed, and for a wide range of portable devices, power. Applications requiring encrypted data must simultaneously optimize the block cipher algorithm and implementation choice against these limitations. To this end we investigate block cipher implementations that are optimized for speed and energy efficiency, the primary metrics of devices such as the MSP430 where constrained memory resources nevertheless allow a range of implementation choices. The results set speed and energy efficiency records for the MSP430 device at 132 cycles/byte and 2.18 J/block for AES-128 and 103 cycles/byte and 1.44 J/block for equivalent block and key sizes using the lightweight block cipher SPECK. We provide a comprehensive analysis of size, speed, and energy consumption for 24 different variations of AES and 20 different variations of SPECK, to aid system designers of microcontroller platforms optimize the memory and energy usage of secure applications.
Cardiovascular diseases are the main cause of death worldwide. Atherosclerosis and atrial fibrillation are structural and electrical pathophysiology, respectively, that can lead to acute events such as stroke or myocardial infarction. We used particle-based Monte Carlo methods to simulate X-ray phase imaging of atherosclerotic plaque types IV-VIII in the aorta, iliac, and coronary arteries. We also assessed scar lesion development in radiofrequency catheter ablation treatment of atrial fibrillation by simulating lesions 2, 5, 10, 30, and 60 days post-procedure. For both applications, we found high signal-to-noise and contrast-to-noise ratios in all lesions. These results suggest that X-ray phase imaging is a viable technique for non-invasive quantitative cardiovascular lesion characterization.
Weave-induced skew on printed wiring boards (PWB) for 10+ Gbps SerDes data rates can be very significant. In this paper, we not only investigate weave-induced skew but also look at other sources of skew. We show the weave skew results taken from measurements of three different test boards. Results from a fourth board are presented to examine PWB differential via skew. Measurements from a fifth board are analyzed to determine total channel skew. We propose a budget such that a certain amount of skew can be tolerated with a small increase in channel insertion loss. We then present a case study to project overall performance on PWB yield. We observe a number of anomalies with our test results and suggest additional studies to guard against unpredicted high skew.
Output driver models play a critical role in simultaneous switching noise (SSN) analysis. However, their accuracy must often be compromised with simplicity of implementation for large scale SSN simulations. We present an approach for creating simple, fast, and accurate macromodels of output drivers. To demonstrate their usefulness, simulation results in multi-IO SSN simulations are shown and compared to those obtained from transistor level SPICE libraries.
We present design, simulation, and measurement of a Ka-band (35 GHz) low noise amplifier (LNA) fabricated in a 120 GHz f t /f m a x SiGe BiCMOS technology (IBM 7HP). To our knowledge, this is the first demonstration of a Ka-band LNA in a SiGe technology, representing the first of a set of desired building blocks for integrating a Ka-band transmit and receive (T/R) module in a single chip environment. At 35 GHz, the 3-stage LNA exhibited 15.1 dB gain, -5.9 dBm output compression (P1dB), 9 dBm third order intercept (IP3), and 5.6 dB noise figure at 25.6 mW DC power. Peak gain and bandwidth of the LNA were found to be 19.0 dB and 10.7 GHz respectively at a center frequency of 31.3 GHz.
A 10 GHz, 15 dB gain feedforward amplifier (FFA) is described whose design is primarily aimed at reducing residual phase noise. The two-module hybrid MIC construction incorporates GaAs MMICs and microstrip thin-film interconnects on alumina. For loop balance electronic gain control is achieved using dual gate distributed amplifiers, while phase control is accomplished with mechanical phase shifters and delay lines. Measurements of near carrier noise are made from 1 Hz to 6.4 MHz. In the flicker region the 1 Hz intercept is -110 dBc/Hz, with an approximate corner frequency around 30 KHz at -160 dBc/Hz.
We present design and characterization of a low power single stage X-band low noise amplifier in an AlSb/InAs HEMT integrated circuit technology. Gain, noise, linearity, and phase noise characterization are presented.