The ULTRA (UAH Logging, Trace Recording, and Analysis) instrumentation system provides an accurate and low-cost means of collecting traces of Message Passing Interface (MPI) program execution. These traces preserve the original parallel program’s data-dependencies by recording each MPI operation performed; the message source, destination, and size; and the number of application instructions preceding the operation. Using these traces enables low-cost simulations for determining the performance of various node and communication hardware combinations although the cluster used to collect the data and the cluster being simulated differ significantly. The portability of ULTRA traces is in direct contrast to time-based measurements which have previously been used. With time-based measurements, it is impossible to remove artifacts caused by network contention or data dependencies, thereby rendering them ineffective for drawing conclusions about any configurations other than the one used to collect the measurements. Further, ULTRA modifications in the Linux kernel allow the Intel Pentium II processor’s performance monitoring hardware to track the application instructions counts on a per thread basis without interference. Thus, multiple threads mapped to single processor in a small cluster can be used to collect accurate trace data for larger clusters. Traces for the NAS BT, SP, CG, EP, IS, LU, and MG benchmarks, representing 5.8x1012 instructions, have been used to drive simulations of PC clusters with fast Ethernet switches and hubs. There was good agreement between the predictions made by the simulators and the original data collection computer. Two representative benchmarks, CG and SP have been selected from the suite and discussed. The portability of the traces is demonstrated by using the data collected for the CG benchmark executing on a cluster with a fast Ethernet switch to accurately predict the performance of a cluster with a fast Ethernet hub. This portability of the trace data allows estimates to be made for networking systems such as Gigabit Ethernet, Myrinet, and proposed optical interconnects.
Just a few years ago, parallel computers were tightly-coupled SIMD, VLIW, or MIMD machines. Now, they are clusters of workstations connected by communication networks yielding ever-higher bandwidth (e.g., Ethernet, FDDI, HiPPI, ATM). For these clusters, compiler research is centered on techniques for hiding huge synchronization and communication latencies, etc. — in general, trying to make parallel programs based on fine-grain aggregate operations fit an existing network execution model that is optimized for point-to-point block transfers. In contrast, we suggest that the network execution model can and should be altered to more directly support fine-grain aggregate operations. By augmenting workstation hardware with a simple barrier mechanism (PAPERS: Purdue's Adapter for Parallel Execution and Rapid Synchronization), and appropriate operating system hooks for its direct use from user processes, the user is given a variety of efficient aggregate operations and the compiler is provided with a more static (i.e., more predictable), lower-latency, target execution model. This paper centers on compiler techniques that use this new target model to achieve more efficient parallel execution: first, techniques that statically schedule aggregate operations across processors, second, techniques that implement SIMD and VLIW execution.
The UAH Logging, Trace Recording, and Analysis instrumentation (ULTRA) provides highly repeatable (0.0002% variation) application instruction counts for parallel programs which are invariant to the communication network used, the number of processors used, and the MPI communication library used. ULTRA, implemented as an MPI profiling wrapper, avoids the data collection system artifacts of time-based measurements by using instruction counts as the basic measure of work performed and records the operation performed and the amount of data sent for each network operation. These measurements can be scaled appropriately for various target architectures. ULTRA's instrumentation overhead is minimized by using the Pentium II processors's performance monitoring hardware, allowing large, production-run applications to be quickly characterized. Traces of the NAS benchmarks representing 6.67×1012 application instructions were generated by ULTRA. The application instructions executed per byte injected into the network and the instructions executed per message sent were computed from the traces. These values can be scaled by the expected processor performance to estimate the minimum network performance required to support the programs. It is impossible to use time-based measurements for this purpose due to measurement artifacts caused by the background processes and the communication network of the data collection system.
Summary form only given, as follows. Mechanisms of microwave pulse shortening are being investigated with a multi-megawatt, large-orbit, coaxial gyrotron. Experiments are primarily concerned with plasma production inside the microwave cavity and electron-beam collector. This gyrotron operates in the S-band at /spl sim/10 MW and is driven by the Michigan Electron Long Beam Accelerator (MELBA) at parameters: V=-750 kV, I/sub beam/=6 kA, I/sub tube/=0.8 kA, and pulselengths of 0.5-2.0 /spl mu/s. Plasma H-/spl alpha/ line radiation is measured inside the microwave cavity and e-beam collector via fiber optic probes/monochromators and correlated with microwave power and microwave cutoff. The temporal correlation is being measured of microwave power to growing H-/spl alpha/ optical emission. A strong correlation would suggest that the plasma is cutting off the microwaves as the plasma reaches critical density (/spl sim/8/spl times/10/sup 10/ cm/sup -3/). Time frequency analysis is used to analyze heterodyne microwave data to investigate other pulse shortening mechanisms, including voltage fluctuations, mode hopping, and mode competition. RF plasma processing is being examined on the coaxial cavity and e-beam collector to determine its effect on the pulse shortening characteristics of this gyrotron device.
Summary form only given. Experiments have proven that both the surface contaminants and microstructure topography on the cathode of an electron beam diode influence impedance collapse and electron emission current. Experiments have characterized effective RF plasma processing protocols for high voltage A-K gaps using argon and argon/oxygen gas mixtures. RF processing time, feed gas pressure, and RF power were adjusted. Time resolved optical emission spectroscopy measured contaminant (hydrogen) and bulk cathode (aluminum) plasma emission versus transported axial electron beam current. Experiments utilize the Michigan Electron Long Beam Accelerator (MELBA) at parameters: V=-0.7 to -1.0 MV, I(diode)=3-30 kA, and pulselength=0.4 to 1.0 microseconds. Microscopic and macroscopic E-fields on the cathode were varied to characterize the scaling of breakdown conditions for contaminants versus the bulk material of the cathode after plasma processing. Electron emission was suppressed for an aluminum cathode in a high voltage A-K gap after RF plasma processing. Experiments using a two-stage low power (100 W) argon/oxygen RF discharge followed by a higher power (200 W) pure argon RF discharge yielded an increase in turn-on voltage required for axial current emission from 662/spl plusmn/174 kV to 981/spl plusmn/97 kV. After two-stage RF plasma processing axial current emission turn-on time was increased from 100/spl plusmn/22 nanoseconds to 175/spl plusmn/42 nanoseconds. Aluminum optical emission was delayed >150 nanoseconds after the overshoot in voltage after two-stage RF plasma processing.
The UAH Logging, Trace Recording, and Analysis instrumentation (ULTRA) provides highly repeatable (0.0002% variation) application instruction counts for parallel programs which are invariant to the communication network used, the number of processors used, and the MPI communication library used. ULTRA, implemented as an MPI profiling wrapper, avoids the data collection system artifacts of time- based measurements by using instruction counts as the basic measure of work performed and records the operation performed and the amount of data sent for each network operation. These measurements can be scaled appropriately for various target architectures. ULTRA's instrumentation overhead is minimized by using the Pentium II processors's performance monitoring hardware, allowing large, production-run applications to be quickly characterized. Traces of the NAS benchmarks representing 6.67x1012 application instructions were generated by ULTRA. The application instructions executed per byte injected into the network and the instructions executed per message sent were computed from the traces. These values can be scaled by the expected processor performance to estimate the minimum network performance required to support the programs. It is impossible to use time-based measurements for this purpose due to measurement artifacts caused by the background processes and the communication network of the data collection system.
Summary form only given, as follows. Large orbit, coaxial gyrotrons are currently under investigation with microwave power up to 40 MW. The electron beam is produced by MELBA (Michigan Electron Beam Accelerator) with the following parameters: 0.75-1.0 MV, 1-10 kA diode current, 0.2-1.5 kA tube current, 0.5-1.0 microsecond pulselength. The axis encircling electron beam is generated by a cusp magnetic field. Electron beam current transport through the magnetic cusp and through the microwave cavity are measured. The frequency of the fundamental coaxial mode, TE/sub 111/, is observed to be 2.34 GHz. A novel method of time-frequency analysis (TF) using reduced interference distributions is used to analyze heterodyne mixer data, TF analysis shows microwave frequency modulation with e-beam voltage modulation. Mode competition and mode hopping between TE/sub 111/ and TE/sub 112/ modes are also observed by TF analysis. Microwave cold test data are compared to operating modes of the gyrotron.
Summary form only given, as follows. Experiments and diagnostics are currently under investigation on large-orbit axis-encircling, coaxial gyrotrons. The electron beam is generated by MELBA (Michigan Electron Long Beam Accelerator) with the following operating parameters: voltage=750-800 kV, diode current=1-10 kA, pulselength =0.5 to 1.0 /spl mu/s. The gyrotron has a coaxial cavity with an inner radius of 3 mm and an outer radius of 3.7 cm. Initial operation of the gyrotron is in the TE11 mode at a frequency of 2.3 GHz. Frequencies of this gyrotron oscillator during operation am measured by Fourier transforming heterodyne mixer data. This high power emission frequency is in agreement with measured cold test resonances and corresponds to calculations for waveguide cutoff of the TE11 mode. Microwave power in excess of 8 MW is measured from the gyrotron oscillator. Microwave power is measured by calibrated crystal detectors. The axis encircling beam is generated by passing the electron beam through a magnetic cusp field. Diagnostic experiments to measure beam alpha (perpendicular velocity/parallel velocity) are conducted using radiation darkened patterns on glass witness plates. Beam alphas of 0.7-1.5 are observed.
Summary form only given. Experiments have proven that both the surface contaminants and the surface topography on the cathode of an e-beam diode influence impedance collapse and emission current. The primary surface contaminant on systems that open to air is H/sub 2/O. Time-resolved optical emission spectroscopy is being used to view contaminant and bulk cathode plasma emission versus transported axial beam current. Experiments utilize the Michigan Electron Long Beam Accelerator (MELBA) at parameters: V=-0.7 to -1.0 MV, I/sub diode/=1-10 kA, and /spl tau//sub e-beam/=0.4 to 1.0 /spl mu/s MELBA is used to study thermal and stimulated desorption of contaminants from anode surfaces due to electron deposition, and breakdown of contaminants from cathode surfaces during the high voltage pulse. Experiments are also underway to characterize effective cleaning protocols for high voltage A-K gaps. RF cleaning techniques using Ar and Ar/O/sub 2/ mixtures are being investigated.
Summary form only given, as follows. Contaminants have been shown to be a contributing factor to parasitic losses in ion beam diodes. Probable primary contaminants are CO, CO/sub 2/, H/sub 2/O, H/sub 2/, and C/sub n/H/sub m/ complexes. Experiments are underway at the University of Michigan to characterize effective cleaning protocols for high voltage A-K gaps. RF cleaning techniques using Ar and Ar/O/sub 2/ mixtures are being investigated. Optical emission spectroscopy is being used to view the effects of cleaning on neutral and ion contaminants during the high voltage diode pulse and to characterize the cleaning discharge. Experiments utilize the Michigan electron long beam accelerator (MELBA) at parameters: V=-0.7 to -1.0 MV, I/sub diode/=1-10 kA, and /spl tau//sub e-beam/=0.4-1.0 /spl mu/s. MELBA is used to study thermal and stimulated desorption of contaminants from A-K surfaces due to electron deposition. Pre-analysis and post-analysis of contaminants (e.g. CO, CO/sub 2/, H/sub 2/, H/sub 2/, C/sub n/H/sub m/) is performed using a residual gas analyzer.
Low latency, high bandwidth interconnection networks that directly link arbitrary pairs of processing elements without contention are very desirable for parallel computers. Most communication networks in parallel machines have made compromises due to the limitations of electronics. Many of the optical interconnection schemes proposed have simply replaced the point-to-point copper wiring with fiber optics and have not made use of the unique properties of optics. This paper proposes an optical interconnect architecture for over a hundred processors, which contains a dedicated channel for each processor to eliminate global arbitration and to provide bandwidth that scales with the number of processors in the machine. Unlike electrical buses, this architecture is not limited by the medium (fiber optics) used to connect the transmitters and receivers. Each processor has an array of receivers, one receiver for each processor channel. The architecture of the receiver array permits a variety of different parallel programming models to be efficiently supported.
Parallel progrm often exhibit strong preferences for different system structrues, and machines with the ideal structures may all be available within a single heterogeneous network.