Modern system-on-chip designs often require multiple clock frequencies. On the other hand, global interconnects suffer large delays. This paper proposes a method that manages these two problems within the framework of conventional synchronous design flow. The design is partitioned into isochronous blocks already at behavioral level, where each block is synchronous using a local clock. The local clock frequencies are assumed related by rational numbers. Communication between blocks is managed with FIFOs at each receiver, which manage different clock frequencies and hide unknown delays or clock skews. This method guarantees clock true implementation of a clock true behavioral description utilizing a predefined block-to-block latency.
Embedded memories in an application specific integrated circuit (ASIC) consume most of the chip area. Data variables of different widths require more memory than needed because they are rounded up to nearest power of 2, i.e., 6 to 8 bits, 11 to 16 bits, and 25 to 32 bits. This can be avoided by adding two bit oriented load and store instructions. The memories can still be 8, 16 or 32 bits wide, but the loads and stores can have arbitrary variable sizes. The hardware changes within the processor are small and an extra hardware block between the processor and the memory is added.
A method to mitigate timing problems due to global wire delays is proposed. The method follows closely a fully synchronous design flow and utilizes only true digital library elements. The design is partitioned into isochronous blocks at system level, where a few clock cycles latency is inserted between the isochronous blocks. This latency is then utilized to automatically mitigate unknown global wire delays, unknown global clock skews and other timing uncertainties occurring in backend design. The new method is expected to considerably reduce the timing closure effort in large high frequency digital designs in deep submicron technologies.
The autocorrelation spectrometer is an important instrument for radio astronomy. In satellite-based spectrometers, low power consumption is essential. The correlator chip presented in this paper reduces the power consumption more than five times compared to other full-custom designs. This has been achieved by reducing the number of clocked transistors, using a compact layout of cells, which reduces wire lengths, and using parallel processing of data. Also, the low power performance is combined with a large number of lags and a high data throughput. The correlator performs 0.5-TMAC operations in 416 lags at a sample rate of 1.28-GSample/s with an input data precision of 1.5-b and a correlation period of one second. The chip is also designed to reduce noise generation by using multiple internal clock phases.
A new interconnect-driven DFT implementation is proposed in this paper. The normal way to implement the DFT is to use the FFT algorithm since it is computationally favorable. However, the increased speed comes at the cost of increased communications which give a higher power consumption. If the DFT algorithm is directly implemented instead, each channel becomes independent of all other channels and consequently communications and hence power consumption are reduced. Other benefits of using the DFT directly are the possibility to calculate a spectrum of any length, not only a power of two, and to have an irregular frequency step between channels. A number of ad hoc processing-element (PE) and system-level solutions are also proposed to reduce the power consumption even further.
Basic limitations to high data throughput chips in CMOS are described and methods for coping with these discussed. The proposed methods are demonstrated by two design examples;: a pipelined datapath architecture for high throughput protocol processing; and a shared buffer architecture for switching.
In telecommunications systems, the commonly used method to generate clocks is based on phase-locked loop or delay-locked loop related frequency synthesis. In this paper, we address a method of digital multiphase clock/pattern generation (MPCG) to generate a system clock or pulse pattern vector when a multiphase clock is available. The advantages of the multiphase clock method are a) the design method is digital; b) the working frequency range is very wide; and c) the sensitivity to noise is less than analog methods. Different approaches to implement the basic blocks in MPCG are described, A design example implemented in BiCMOS uses eight clock phases at 622 MHz obtained by dividing a 5-GHz clock to generate a clock at 622 MHz x 32/53 = 376 MHz. By such a method, we can generate a pulse pattern vector as well. The maximum time resolution is equal to half of the phase difference. A low power solution is achieved without loss of circuit speed.
This quadrature DDFS calculates sine and cosine values with a tuning resolution below 1 Hz, by only using an 8 word ROM and interpolation. Two internal 8-bit differential D/A converters generate the four-phase analog output signal. A spurious free dynamic range of 50 dB for low frequencies and 30 dB near Nyquist is achieved.
The requirement for higher bandwidth in the fixed network makes 10 Gb/s working in the SDH/SONET format an attractive alternative. Circuits performing the processing of SDH frames are needed. The framer is a part of a regenerator, the simplest SDH equipment. In addition to reshaping, the regenerator performs retiming and regenerating (3R) measurements on the status of the connection such as error rate, and communicates that information back to the operation and maintenance system. Multiplexers on the 2B level (l6:l) have a complexity that makes possible implementation in commercial technology. This gives a serial speed of 622 Mb/s at the framer-interface. CMOS has the power to operate at this I/O speed and with a 622 MHz internal clock. This facilitates a compact and cost-effective solution. Direct interfacing to STM-4/OC-12 (622 Mb/s) is also possible. The high clock frequency leads to a high degree of pipelining. The extra pipelining does not significantly affect the performance of the system since each pipeline only delays the signal as much as one extra foot of fibre.
Today, it is accepted to say that the more bits you have got in your micro processor the better performance you will have. This is probably true if only the performance is concerned. However, if the chip size of the processor is taken into account this might not be the case. In massively parallel architecture, chip area is an important figure. This is especially true for air-borne and, to certain extend, industrial real-time applications. In this paper we study the impact of multi bit processors on the linear SIMD array called RVIP which is an architecture used for real-time radar signal processing. In RVIP, each processing element is bit serial. The results show that the gain in using multi bit processors is very little and that the optimal bit number probably is one or two. We believe that the results from this study can be transferred to other similar systems.
We show that the linear SIMD architecture FVIP is capable of performing a typical MPD radar signal processing including FFT and the resolving algorithm. A conventional DSP system clocked at 50 MHz would require approximately 200 DSPs to have the same performance as our FVIP system. We have estimated the size and power consumption of our FVIP system to be 3 dm3 and 200 W. With these MPD studies together with previous LPD studies we are confident that a “VIP-type” architecture can handle all typical pulse Doppler radar waveforms currently used in airborne radar