Experimental results are shown in Tables 1 and 2. We compare tree length (TL), the sum total path length for all p i ! p j pairs (TPL), and tree diameter D. Also included are the maximum delay between any source-sink pair (MD) and the average of maximum delays for each source (AMD), averaged over all runs. Delays were measured from the input transition to the output reaching 90% of its nal value. Delays are shown in nanoseconds, while tree and path lengths are shown in centimeters. Table 3 shows the solutions by 1-Steiner, MD A-tree, and MC MD A-tree when only a (randomly chosen) subset of pairs are critical. This table assumes the CMOS IC parameters, and 8 points per test set. We report the total weighted path length TWPL, the required weighted path length RWPL (a lower bound), and the maximum and average maximum delays for the critical pairs. Compared with the best known Steiner tree heuristic, minimum cost minimum diameter A-trees ooer maximum delay improvements of 1% to 11% and 4% to 13% for 0:5 CMOS ICs and MCMs. Average maximum delays were improved by 0% to 4% and 1% to 12%. When only a subset of point pairs are critical, maximum delay improvements for 0:5 CMOS IC routings were from 3% to 7%. Table 3: Comparison of routing results for test cases where only a limited number of point pairs have W(p i ; p j) = 1. Figure 1: The smallest tilted rectangle containing the points, and a smallest tilted square which contains the rectangle. In fact, it is not always necessary to place the root of the A-tree at the center of an STS. It is easy to see that as long as the root satisses the constraint d(p i ; r) + d(r; p j) D for all p i and p j , the tree will have minimum diameter. In the Euclidean plane, the region which satisses the constraint is simply the intersection of all ellipses formed by pairs p i , p j. We deene an octilinear segment to be a segment that is either horizontal, vertical, or have slope 1. In the Manhattan plane, the set of points satisfying d(p i ; r) + d(r; p j) D is an octilinear ellipse (OE), bounded by no more than eight octilinear segments. If d(p i ; p j) = D 0 …
The authors model high-speed VLSI interconnects by using a generic distributed RLC-tree. Through a detailed analysis of the distributed RLC-tree a two-pole approximation system is established to formulate the performance-driven layout in MCM designs. An in-depth study of the formulated performance-driven layout problem reveals the interplay between the interconnector's performance and its geometrical parameters. The study leads to an A-tree topology to optimize the defined performance-driven layout problem. Significant improvement, an average up to 67% reduction on the interconnection delay, is achieved over large sample MCM designs, as compared with the well known Steiner tree topology.<>
those of the MST and AHHK 1] constructions for nets of up to 17 pins using the same IC parameters. All delays in the table are calculated using the Two-Pole simulator. The AHHK algorithm of Alpert et al. is a recent cost-radius tradeoo construction which yields less tree cost (and signal delay) for given tree radius bounds than the method of 3]. Our results indicate that the LDT algorithm is highly eeective for larger nets, and also outperforms the best known direct trade-oo between tree radius and cost (i.e., AHHK): for 17-pin nets, the LDT construction reduces average sink delay by 35.4% compared to MSTs and by 6:2% compared to AHHK trees. Table 5: Two-Pole simulator delays comparing LDT with MST and AHHK trees on nets with up to 17 pins. Averages in each column are taken over 500 random nets. 6 Conclusions We have addressed the issue of the accuracy and delity of the Elmore 7] and Two-Pole 17] delay models by comparing their rankings of tree topologies with rankings by the SPICE3e2 simulator. Our studies indicate that algorithms which minimize the Elmore and Two-Pole delay estimates should also eeectively minimize actual delay. We have also described a branch-and-bound BBORT method to determine optimal routing trees for any given monotone delay function. To achieve a practical and near-optimal routing methodology, we have proposed the greedy Low Delay Tree (LDT) heuristic. LDT can be implemented using any given model of delay; because of the demonstrated delity of Elmore delay, we have implemented LDT using that model. Experimental results show that LDT performs essentially as well as exhaustive search on nets with up to 7 pins. In addition, LDT achieves reductions in delay of up to 35% (depending on the net size) over the MST routing, as measured by the Two-Pole simulator. Signiicant reductions are also achieved compared to AHHK-1] and SPT-based routings. LDT is formulated to construct a spanning tree, but can easily be extended to yield a Steiner Low Delay Tree (SLDT) algorithm. For example, we may allow each newly selected pin to connect to the closest point in any existing tree edge, possibly inducing a Steiner point. Simulation results in 2] indicate that the SLDT algorithm using Elmore delay is also highly eeective. LDT can also be generalized to \critical-sink routing" by modifying the objective function in the LDT and SLDT algorithms to minimize delay at prescribed critical …
The delays at all sink nodes were measured using the two-pole circuit sim-ulator proposed by Zhou et al. 17] and discussed in 3]. This simulator is a computationally eecient code which has produced very accurate results (within a few percent) when tested against SPICE. We consider both the average delay (taken over all sinks) and the worst-case delay (i.e., the latest arrival time of the signal to any sink); all results are normalized to the corresponding values for the MST routing. We make the following observations. 1. While BRBC tends to yield lower tree cost for any xed , if we consider the family of trees over each instance, we see that the AHHK solution will almost always have lower cost for any given tree radius. 2. AHHK yields lower worst-case and average-case signal delay than BRBC (Table 2). As nets become larger, MST radius becomes quite large, and the delay improvements from the AHHK solution become more obvious. 3. While tree cost closely reeects Elmore delay in the IC technology, diierences between the algorithms are small since both will closely approximate the MST for those values of that minimize delay. On the other hand, since tree radius is dominant for MCM interconnects, AHHK is superior in this technology since it gives less cost for a given radius. 4. For each instance of each net size, we recorded the c and values that aaorded lowest signal delay over the entire family of 21 trees generated by each algorithm. The average \best" parame-terization is reported in Table 3, and is useful in guiding the application of AHHK when it is too expensive to generate an entire family of trees. 5 Conclusions Analysis of distributed RC delay shows that performance-driven routing requires tree constructions that can trade oo cost and radius according to interconnect technology and net size. Previous approaches 2, 4, 10] essentially rely on a depth rst traversal of the MST and insert shortest paths as needed to maintain a prescribed radius bound. In contrast, our new AHHK approach directly combines the recurrences for Prim's MST algorithm and Dijk-stra's SPT algorithm. The result is an elegant trade-oo between radius and cost which empirically yields lower-cost trees than the BRBC algorithm 4] for any given tree radius. Simulation results show that the AHHK algorithm constructs routing trees with sig-niicantly less maximum and average delay than the BRBC method 4], in both …
Charles J. Alpert合作论文数IBM Austin Research Laboratory;IBM Research Division3