The traditional SerDes link simulation process begins with the extraction of printed circuit board (PCB) physical stripline and via models, followed by channel modeling and link simulation. We invert this simulation flow by first creating link performance curves across an array of hypothetical channels defined with specially-developed, high level, equation-based models; limited physical extraction is later undertaken to relate PCB channel implementation to these performance curves. These curves allow us to determine the system-level SerDes channel requirements and to become better informed in choosing PCB technologies for lower cost and easier manufacturability. The inverted modeling process is very efficient, allowing for the rapid identification and avoidance of problematic channel topologies and the study of other potentially useful channel designs.
Variability analysis is important in successfully deploying multi-gigabit backplane printed wiring boards (PWBs) with growing numbers of high-speed SerDes links. We discuss the need for large sample sizes to obtain accurate variability estimates of SI metrics (eye height, phase skew, etc). Using a dataset of 11,961 S-parameters, we demonstrate statistical techniques to extract accurate estimates of PWB SI performance variations. We cite numerical examples illustrating how these variations may contribute to underestimated or overestimated design criteria, causing unnecessary design expense. Tabular summaries of performance variation and key findings of broad interest to the general SI community are highlighted.
This paper presents a power distribution network (PDN) decoupling capacitor optimization application with three primary goals: reduction of solution times for large networks, development of flexible network scoring routines, and a concentration strictly on achieving the best network performance. Example optimizations are performed using broadband models of a printed circuit board (PCB), a chip-package, on-die networks, and candidate capacitors. A novel worst-case time-domain optimization technique is presented as an alternative to the traditional frequency-domain approach. The trade-offs and criteria for scoring the computed network are presented. The output is a recommended set of capacitors which can then be applied to the product design.
High performance computing (HPC) systems make extensive use of high speed electrical interconnects, in routing signals among processing elements, or between processing elements and memory. Increasing bandwidth demands result in high density, parallel I/O exposed to crosstalk due to tightly coupled transmission lines. The crosstalk cancellation signaling concept discussed in this paper utilizes the known, predictable theory of coupled transmission lines to cancel crosstalk from neighboring traces with carefully chosen resistive cross-terminations between them. Through simulation and analysis of practical bus architectures, we explore the merits of crosstalk cancellation which could be used in dense interconnect HPC (or other) applications.
Complex digital systems such as high performance computers (HPCs) make extensive use of high-speed electrical interconnects, in routing signals among processing elements, or between processing elements and memory. Despite increases in serializer/deserializer (SerDes) and memory interface speeds, there is demand for higher bandwidth busses in constrained physical spaces which still mitigate simultaneous switching noise (SSN). The concept of zero sum signaling utilizes coding across a data bus to allow the use of single-ended buffers while still mitigating SSN, thereby reducing the number of physical channels (e.g. circuit board traces) by nearly a factor of two when compared with traditional differential signaling. Through simulation and analysis of practical (non-ideal) data bus and power delivery network architectures, we demonstrate the feasibility of zero sum signaling and compare performance with that of traditional (single-ended and differential) methods.
Although next generation (>28 Gbps) SerDes standards have been contemplated for several years, it has not been clear whether PCB structures supporting 56 Gbps NRZ will be feasible and practical. In this paper, we assess a number of specific PCB design strategies (related to pin-field breakouts, via stubs, and fiber weave skew) both through simulation and through measurement of a wide range of structures on a PCB test vehicle. We demonstrate that conventional approaches in many cases will not be sufficient, but that modest (manufacturable) design changes can enable low-skew 56 Gbps NRZ channels having acceptable insertion and return loss.
Evaluating and scaling direct liquid cooling (DLC) hardware from the component-level to system-level can present a unique set of challenges. Component-level thermal and flow performance specifications are not always readily available or may be generated under a different set of flow conditions (e.g., flow regime, fluid type) by the manufacturer that may not align with the design engineer’s application space. In addition, increasing component heat dissipation levels are driving the development of high-performance micro-channel cold plates that create an additional pressure drop burden on supporting cooling infrastructure (e.g., cooling distribution units and facility chillers) [1]. Establishing accurate thermal and flow performance characteristics at the component-level early in the DLC design cycle is imperative when building a foundation to predict realistic flow requirements that scale from the component-to-system levels, while also including design considerations for facility infrastructure performance constraints. Mayo Clinic SPPDG designed and assembled a pressure/flow measurement system and developed a more accurate measurement technique, which incorporates a de-embedding methodology. This paper describes the test measurement system, referred to as the Thermal Test Cart; defines the de-embedded measurement technique, comparing it with conventional parallel measurement processes; and demonstrates the significance of this methodology when applied at the system level.
An earlier study of a high layer-count test board using plated-through-hole (PTH) vias and a limited quantity of laser vias was shown to be capable of supporting 112 Gb/s PAM-4 links (or equivalent signaling having 28 GHz (Nyquist) bandwidth). This original board design was then rebuilt using a different fabricator, and the test results revealed a significant decrease in the bandwidth of the vias. These results led to the development of a set of design specifications that PCB vendors can easily validate, which will ensure that the use of high layer-count boards with PTH technology are viable for emerging 112 Gb/s PAM-4 links.
Signal integrity analysis often involves the development of design guidelines through manual manipulation of circuit parameters and judicious interpretation of results. Such an approach can result in significant effort and sub-optimal conclusions. Optimization routines have been well proven to aid analysis across a variety of common tasks. In addition, there are several non-traditional applications where optimization can be useful. This paper begins by describing the basics of optimization followed by two specific case studies where non-traditional optimization provides significant improvements in both analysis efficiency and channel performance.
Coding schemes are often used in high-speed processor-processor or processor-memory busses in digital systems. In particular, we have introduced (in a 2012 DesignCon paper) a zero sum (ZS) signaling method which uses balanced or nearly-balanced coding to reduce simultaneous switching noise (SSN) in a single-ended bus to a level comparable to that of differential signaling. While several balanced coding schemes are known, few papers exist that describe the necessary digital hardware implementations of (known) balanced coding schemes, and no algorithms had previously been developed for nearly-balanced coding. In this work, we extend a known balanced coding scheme to accommodate nearly-balanced coding and demonstrate a range of coding and decoding circuits through synthesis in 65 nm CMOS. These hardware implementations have minimal impact on the energy efficiency and area when compared to current serializer/deserializers (SerDes) at clock rates which would support SerDes integration.
Significance: A path is described to increase the sensitivity and accuracy of body-worn devices used to monitor patient health. This path supports improved health management. A wavelength-choice algorithm developed at Mayo demonstrates that critical biochemical analytes can be assessed using accurate optical absorption curves over a wide range of wavelengths. Aim: Combine the requirements for monitoring cardio/electrical, movement, activity, gait, tremor, and critical biochemical analytes including hemoglobin makeup in the context of body-worn sensors. Use the data needed to characterize clinically important analytes in blood samples to drive instrument requirements. Approach: Using data and knowledge gained over previously separate research threads, some providing currently usable results from more than eighty years back, determine analyte characteristics needed to design sensitive and accurate multiuse measurement and recording units. Results: Strategies for wavelength selection are detailed. Fine-grained, broad-spectrum measurement of multiple analytes transmission, absorption, and anisotropic scattering are needed. Post-Beer-Lambert, using the propagation of error from small variations, and utility functions that include costs and systemic error sources, improved measurements can be performed. Conclusions: The Mayo Double-Integrating Sphere Spectrophotometer (referred hereafter as MDISS), as described in the companion report arXiv:2212.08763, produces the data necessary for optimal component choice. These data can provide for robust enhancement of the sensitivity, cost, and accuracy of body-worn medical sensors. Keywords: Bio-Analyte, Spectrophotometry, Body-worn monitor, Propagation of error, Double-Integrating Sphere, Mt. Everest medical measurements, O2SAT Please see also arXiv:2212.08763
As ASIC supply voltages approach one volt, the source-impedance goals for power distribution networks are driven ever lower as well. One approach to achieving these goals is to add decoupling capacitors of various values until the desired impedance profile is obtained. An unintended consequence of this approach can be reduced power supply stability and even oscillation. In this paper, we present a case study of a system design which encountered these problems and we describe how these problems were resolved. Time-domain and frequency-domain analysis techniques are discussed and measured data is presented.
It is notoriously difficult to measure instantaneous supply current to a device such as an ASIC, FPGA, or CPU without also affecting the instantaneous supply voltage and compromising the operation of the device [21]. For decades designers have relied on rough estimates of dynamic load currents that stimulate a designed Power Delivery Network (PDN). The consequences of inaccurate load-current characterization can range from excessive PDN cost and lengthened development schedules to poor performance or functional failure. This paper will introduce and describe a method to precisely determine timedomain current waveforms from a pair of measured timedomain voltage waveforms. This NonInvasive Current Estimation (NICE) method is based on established twoport network theory along with component and board modeling techniques that have been validated through measurements on demonstrative circuits. This paper will show that the NICE method works for any transient event that can be captured on a digital oscilloscope. Limitations of the method and underlying measurements are noted where appropriate. The method is applied to a simple PDN with an arbitrary load, and the NICE-derived current waveform is verified against an independent measurement by sense resistor. With careful component and board modeling, it is possible to calculate current waveforms with a root mean square error of less than five percent compared to the reference measurement. Current transients that were previously difficult or impossible to characterize by any means can now be calculated and displayed within seconds of an oscilloscope-trigger event by using NICE. ASIC and FPGA manufacturers can now compute the startup current for their device and publish the actual waveform, or provide a piecewiselinear SPICE model (PWL source) to facilitate design and testing of the regulator and PDN required to support their device.
The advantages and limitations of time-domain pseudo-random binary sequence (PRBS) excitation methods for system identification of individual modes within a multi-conductor transmission system are discussed. We develop the modifications necessary to standard frequency-domain transmission-line models to match time-domain experimental data from several types of transmission systems. We show a variety of experimental results showing very good to excellent agreement with our model's predictions, up to approximately 10 GHz.
As data rates for multi-gigabit serial interfaces within multi-node compute systems approach and exceed 10 Gigabits per second (Gbps), board-to-board and chip-to-chip optical signaling solutions become more attractive, particularly for longer (e.g. 50-100 cm) links. The transition to optical signaling will potentially allow new high performance compute (HPC) system architectures that benefit from characteristics unique to optical links. To examine these characteristics, we built and tested several optical demonstration vehicles; one based on dense wavelength division multiplexing (DWDM), and others based on multiple point-to-point links carried across multimode fibers. All test vehicles were constructed to evaluate applicability to a multi-node compute system. Test results, combined with data from recent research efforts are summarized and compared to equivalent electrical links and the advantages and design characteristics unique to optical signaling are identified.
Power delivery network (PDN) model development is often simplified using superports or pin-groups for high pin count devices. This approach significantly reduces model complexity but can compromise accuracy in holistic time- and frequency-domain analyses. In the context of the non-invasive current estimation (NICE) technique for packaged, high-performance integrated circuits (ICs), this paper describes the limitations of using pin-group PDN models in distributed impedance applications. DC error calibration factors are calculated, and tuned AC error compensation networks are proposed. These fundamental techniques for working with overly simplified models in highly distributed power integrity applications are derived and demonstrated through simulation and measurement on exemplar hardware.
Applications such as public-key cryptography are critically reliant on the speed of modular multiplication for their performance. This paper introduces a new block-based variant of Montgomery multiplication, the Block Product Scanning (BPS) method, which is particularly efficient using new 512-bit advanced vector instructions (AVX-512) on modern Intel processor families. Our parallel-multiplication approach also allows for squaring and sub-quadratic Karatsuba enhancements. We demonstrate $$1.9\,\times $$ improvement in decryption throughput in comparison with OpenSSL and $$1.5\,\times $$ improvement in modular exponentiation throughput compared to GMP-6.1.2 on an Intel Xeon CPU. In addition, we show $$1.4\,\times $$ improvement in decryption throughput in comparison with state-of-the-art vector implementations on many-core Knights Landing Xeon Phi hardware. Finally, we show how interleaving Chinese remainder theorem-based RSA calculations within our parallel BPS technique halves decryption latency while providing protection against fault-injection attacks.
Techniques for predicting or measuring the dynamic current of a high-performance integrated circuit have proven useful when applied in theory. However, the practical application of such methodologies in the lab is much more challenging. Test equipment imperfections have significant influence on the dynamic and static voltage measurement accuracy and thereby impede accurate prediction of chip load current. This paper empirically addresses the presence of ground loops in the measurement system and demonstrates appropriate isolation techniques to resolve the two-channel oscilloscope voltage measurements necessary for accurate load current calculation. In addition, these findings are correlated with simulations by adding oscilloscope and cabling parasitic inductance and resistance to simulation models, which further demonstrate the importance of ground isolation.
We compare the energy consumed by 8-bit x 8-bit and 16-bit x 16-bit multipliers composed of small analog multipliers implemented using transistors operating in their sub-threshold regions to the energy consumed by equivalent digital implementations. The analysis shows that the analog energy consumption is determined by the required signal-to-noise ratio of the individual analog multipliers and that the energy consumption is higher than the equivalent digital implementation’s energy consumption.
Voltage regulator impedance characterization is a straightforward concept but the challenges in making high-fidelity measurements are numerous as the impedance becomes progressively smaller. This paper evaluates techniques used for low impedance buck converters requiring large-signal current perturbations to properly characterize the impedance. Additionally, the paper addresses how high-performance time-domain characterization equipment can be utilized to generate frequency-domain impedance plots. A number of parameters that can affect the impedance plots were studied and tuned. Some of the measurement challenges when collecting this data are examined.