A high-speed VLSI interconnection is modeled by using a generic distributed-RLC tree. Through a detailed analysis of the distributed-RLC tree a two-pole approximation system is established. Relating the solution of the two-pole approximation to the interconnection geometric parameters, a VLSI performance-driven layout problem is formulated which reveals the interplay between the interconnection performance and its geometrical parameters.<>
Experimental results are shown in Tables 1 and 2. We compare tree length (TL), the sum total path length for all p i ! p j pairs (TPL), and tree diameter D. Also included are the maximum delay between any source-sink pair (MD) and the average of maximum delays for each source (AMD), averaged over all runs. Delays were measured from the input transition to the output reaching 90% of its nal value. Delays are shown in nanoseconds, while tree and path lengths are shown in centimeters. Table 3 shows the solutions by 1-Steiner, MD A-tree, and MC MD A-tree when only a (randomly chosen) subset of pairs are critical. This table assumes the CMOS IC parameters, and 8 points per test set. We report the total weighted path length TWPL, the required weighted path length RWPL (a lower bound), and the maximum and average maximum delays for the critical pairs. Compared with the best known Steiner tree heuristic, minimum cost minimum diameter A-trees ooer maximum delay improvements of 1% to 11% and 4% to 13% for 0:5 CMOS ICs and MCMs. Average maximum delays were improved by 0% to 4% and 1% to 12%. When only a subset of point pairs are critical, maximum delay improvements for 0:5 CMOS IC routings were from 3% to 7%. Table 3: Comparison of routing results for test cases where only a limited number of point pairs have W(p i ; p j) = 1. Figure 1: The smallest tilted rectangle containing the points, and a smallest tilted square which contains the rectangle. In fact, it is not always necessary to place the root of the A-tree at the center of an STS. It is easy to see that as long as the root satisses the constraint d(p i ; r) + d(r; p j) D for all p i and p j , the tree will have minimum diameter. In the Euclidean plane, the region which satisses the constraint is simply the intersection of all ellipses formed by pairs p i , p j. We deene an octilinear segment to be a segment that is either horizontal, vertical, or have slope 1. In the Manhattan plane, the set of points satisfying d(p i ; r) + d(r; p j) D is an octilinear ellipse (OE), bounded by no more than eight octilinear segments. If d(p i ; p j) = D 0 …
The limiting factor for high-performance systems is being set by interconnection delay rather than transistor switching speed. The advances in circuits speed and density are placing increasing demands on the performance of interconnections, for example chip-to-chip interconnection on multichip modules. To address this extremely important and timely research area, we analyze in this paper the circuit property of a generic distributedRLC tree which models interconnections in high-speed IC chips. The presented result can be used to calculate the waveform and delay in anRLC tree. The result on theRLC tree is then extended to the case of a tree consisting of transmission lines. Based on an analytical approach a two-pole circuit approximation is presented to provide a closed form solution. The approximation reveals the relationship between circuit performance and the design parameters which is essential to IC layout designs. A simplified formula is derived to evaluate the performance of VLSI layout.
The delays at all sink nodes were measured using the two-pole circuit sim-ulator proposed by Zhou et al. 17] and discussed in 3]. This simulator is a computationally eecient code which has produced very accurate results (within a few percent) when tested against SPICE. We consider both the average delay (taken over all sinks) and the worst-case delay (i.e., the latest arrival time of the signal to any sink); all results are normalized to the corresponding values for the MST routing. We make the following observations. 1. While BRBC tends to yield lower tree cost for any xed , if we consider the family of trees over each instance, we see that the AHHK solution will almost always have lower cost for any given tree radius. 2. AHHK yields lower worst-case and average-case signal delay than BRBC (Table 2). As nets become larger, MST radius becomes quite large, and the delay improvements from the AHHK solution become more obvious. 3. While tree cost closely reeects Elmore delay in the IC technology, diierences between the algorithms are small since both will closely approximate the MST for those values of that minimize delay. On the other hand, since tree radius is dominant for MCM interconnects, AHHK is superior in this technology since it gives less cost for a given radius. 4. For each instance of each net size, we recorded the c and values that aaorded lowest signal delay over the entire family of 21 trees generated by each algorithm. The average \best" parame-terization is reported in Table 3, and is useful in guiding the application of AHHK when it is too expensive to generate an entire family of trees. 5 Conclusions Analysis of distributed RC delay shows that performance-driven routing requires tree constructions that can trade oo cost and radius according to interconnect technology and net size. Previous approaches 2, 4, 10] essentially rely on a depth rst traversal of the MST and insert shortest paths as needed to maintain a prescribed radius bound. In contrast, our new AHHK approach directly combines the recurrences for Prim's MST algorithm and Dijk-stra's SPT algorithm. The result is an elegant trade-oo between radius and cost which empirically yields lower-cost trees than the BRBC algorithm 4] for any given tree radius. Simulation results show that the AHHK algorithm constructs routing trees with sig-niicantly less maximum and average delay than the BRBC method 4], in both …
Charles J. Alpert合作论文数IBM Austin Research Laboratory;IBM Research Division2