Adaptive optics (AO) is a technique for correcting aberrations introduced when light propagates through a medium, for example, the light from stars propagating through the turbulent atmosphere. The components of an AO instrument are: (1) a camera to record the aberrations, (2) a corrective mechanism to correct them, (3) a real-time controller (RTC) that processes the camera images and steers the corrective mechanism on milliseconds timescales. We have accelerated the image processing for the AO RTC with the use of graphics processing units (GPUs). It is crucial that the image is processed before the atmospheric turbulence has changed, i.e., in one or two milliseconds. The main task is to transfer the images to the GPU memory with a minimum delay. The key result of this paper is a demonstration that this can be done fast enough using commercial frame grabbers and standard CUDA tools. Our benchmarking image consists of \(1.6 \times 10^6\) pixels out of which \(1.2 \times 10^6\) are used in processing. The images are characterized and reduced into a set of 9248 numbers; about one-third of the total processing time is spent on this characterization. This set of numbers is then used to calculate the commands for the corrective system, which takes about two-third of the total time. The processing rate achieved on a single GPU is about 700 frames per second (fps). This increases to 1100 fps (1565 fps) if we use two (four) GPUs. The variation in processing time (jitter) has a root-mean-square value of 20–30 \(\upmu \)s and about one outlier in a million cycles.
Adaptive Optics (AO) is a necessary technology for ensuring the success of the next generation of extremely large telescopes (ELTs). It’s used to help mitigate the perturbing effects of Earth’s atmosphere on the incoming light from astronomical objects and will be an integral part of ELTs for obtaining close to diffraction limited images. To maintain a correction of the incoming wavefront under dynamic atmospheric conditions, which can change significantly on the order of milliseconds, the frame-by-frame reconstruction must be operated in real-time, with hard limits on the time interval between measuring the disturbance and applying a correction. The main problem size for AO RTC increases with the 4th power of telescope diameter and so the computational demands of AO RTCs for ELTs, with primary mirror diameters between 20-40m, increase significantly compared to the current generation of 10m class telescopes. This makes the investigation into and the development of real-time controllers (RTCs) for ELT scale AO systems critical for ensuring the effectiveness of these instruments for first light. Green Flash, which is an ongoing EU funded project, has the aim of investigating the optimal hardware architecture for ELT scale AO RTC, with an emphasis on GPU and Xeon Phi solutions. The Intel Xeon Phi, built using Intel’s Many Integrated Core (MIC) architecture, incorporates ≥64 general purpose x86 CPU cores into a single CPU package paired with a large pool of on chip high bandwidth MCDRAM, it has many of the advantages of current technologies without some of the more significant drawbacks. The most computationally intensive aspects of most AO RTC pipelines are large matrix-vector multiplications mainly used to compute the reconstructed wavefronts which are highly parallelizable and are generally memory bandwidth bound. This makes the Xeon Phi with it’s large CPU count and high bandwidth memory ideally suited for acceleration of the reconstruction task and therefore for ELT scale AO RTC. The most recent incarnation of the Xeon Phi platform is available as a standard socketed x86 CPU allowing previous efforts made in developing CPU based RTC software to be used as a basis for a Xeon Phi based RTCs with the added advantage that any optimisations made for the MIC architecture can be carried forward to future x86 CPU based systems. The Durham Adaptive Optics Real-time Controller (DARC) is an example of a freely available, on-sky tested, fully modular, x86 CPU based AO RTC which which is ideally suited to be a basis for our investigation into ELT scale AO RTC performance. We present a proof of concept AO RTC system, in collaboration with the Green Flash project, using an optimised DARC on a multi-node homogeneous Xeon Phi cluster to demonstrate the potential of the MIC platform for AO RTC. We will present our methods of optimisation for the C based DARC for the Xeon Phi, including BIOS, kernel and OS tuning as well as considerations for multi-threading and massively parallel algorithm development.
With the next-generation of Extremely Large Telescopes (ELTs), the demands of adaptive optics real-time control (AO RTC) increase massively compared to the most complex AO systems in use today. Green Flash, an ongoing EU funded project, is investigating the optimal architecture for ELT scale AO RTC, with an emphasis on GPU and many core CPU solutions. The Intel Xeon Phi range of x86 CPUs is our current focus of investigation into CPU technologies to solve the ELT-scale AO RTC problem. Built using Intels Many Integrated Core (MIC) architecture incorporating 64 general purpose x86 CPU cores into a single CPU package paired with a large pool of on-chip high bandwidth MCDRAM, the Xeon Phi includes many of the advantages of current technologies. The current generation Xeon Phi is readily compatible with standard Linux operating systems and all of the tools and libraries, and as a standard socketed CPU it eliminates the latency introduced by the extra data transfers required for previous Xeon Phis and other accelerator devices. The Durham Adaptive Optics Real-time Controller (DARC) is a freely available, on-sky tested, fully modular, x86 CPU based AO RTC which which is ideally suited to be a basis for our investigation into ELT scale AO RTC performance. We present a proof of concept AO RTC system, in collaboration with the Green Flash project, for ELT scale MCAO, with the requirements of the MAORY AO system in mind, using an optimised DARC on Xeon Phi hardware to achieve the required performance.
The Bateman equations model the evolution of nuclide number densities as a function of time when they are subjected to irradiation. When particles such as neutrons, protons and deuterons interact with a material, several reactions can occur. Depending on the nuclides which are being irradiated, the nuclide may absorb the incoming particle, collide with the incoming particle causing scattering or it may split up into two smaller nuclides as a result of fission. A burn up code models the chains of reactions with respect to time by utilizing nuclear databases that define the cross-sections and decay constants for every nuclide and solves the Bateman eqautions, described in matrix form as Ṅ= AN with N(0)=N 0 . Therefore to solve the Bateman equations, one must compute the matrix exponential N=e At N 0 . There are many methods of calculating the matrix exponential 1 but most are unsuitable for activation matrices found in burn up calculations where there are large differences over several orders of magnitude between values in the activation matrix. However, the Chebyshev rational approximation method (CRAM) provides a robust and accurate solution to burn up equations, treating both short-lived and long-lived nuclides simultaneously and has a very short computation time 2 . Currently in nuclear fusion applications for a reactor model with many cells, burn up/activation solutions for each cell are calculated sequentially. This work presents an accelerated Bateman solver based on CRAM using graphics processing unit (GPU) technology to enable solving a large number of sets of Bateman equations, for many reactor cells simultaneously.
Recent advances in adaptive optics (AO) have led to the implementation of wide field-of-view AO systems. A number of wide-field AO systems are also planned for the forthcoming Extremely Large Telescopes. Such systems have multiple wavefront sensors of different types, and usually multiple deformable mirrors (DMs). Here, we report on our experience integrating cameras and DMs with the real-time control systems of two wide-field AO systems. These are CANARY, which has been operating on-sky since 2010, and DRAGON, which is a laboratory AO real-time demonstrator instrument. We detail the issues and difficulties that arose, along with the solutions we developed. We also provide recommendations for consideration when developing future wide-field AO systems.
The synthetic aperture microwave imaging diagnostic has been operating on the MAST experiment since 2011. It has provided the first 2D images of B-X-O mode conversion windows and showed the feasibility of conducting 2D Doppler back-scattering experiments. The diagnostic heavily relies on field programmable gate arrays to conduct its work. Recent successes and newly gained experience with the diagnostic have led us to modify it. The enhancements will enable pitch angle profile measurements, O and X mode separation, and the continuous acquisition of 2D DBS data. The diagnostic has also been installed on the NSTX-U and is acquiring data since May 2016.
CANARY is an on-sky demonstrator instrument for the investigation of novel forms of tomographic Adaptive Optics (AO) which are required by the prospective European Extremely Large Telescope (E-ELT). CANARY is deployed on the 4.2m William Herschel Telescope in the Canary Islands, Spain. In its most evolved variant (Phase C2), CANARY employs four Laser Guide Star (LGS) Wavefront Sensors (WFS) and three Natural Guide Star (NGS) WFS to perform tomography on the turbulent atmospheric volume above the telescope. The projected wavefront phase error for a particular line of sight is then corrected by two Deformable Mirrors (DMs), and the resultant residual wavefront error is monitored by a further NGS WFS and an image recording camera. We outline the CANARY design, and present the most recent results obtained with a novel open/closed loop control approach designed to mimic the E-ELT.
CANARY is an on-sky Laser Guide Star (LGS) tomographic AO demonstrator in operation at the 4.2m William Herschel Telescope (WHT) in La Palma. From the early demonstration of open-loop tomography on a single deformable mirror using natural guide stars in 2010, CANARY has been progressively upgraded each year to reach its final goal in July 2015. It is now a two-stage system that mimics the future E-ELT: a GLAO-driven woofer based on 4 laser guide stars delivers a ground-layer compensated field to a figure sensor locked tweeter DM, that achieves the final on-axis tomographic compensation. We present the overall system, the control strategy and an overview of its on-sky performance.
An astronomical adaptive optics test-bench, designed to replicate the conditions of a 4 m-class telescope, is presented. Named DRAGON-Next Generation, it is constructed primarily from commercial off-the-shelf components with minimal customization (approximately a 90:10 ratio). This permits an optical design which is modular and this leads to a reconfigurability. DRAGON-NG has been designed for operation for the following modes: (high-order) SCAO, (twin-DM) MOAO, and (twin-DM) MCAO. It is capable of open-loop or closed-loop operation, with (3) natural and (3) laser guide-star emulation at loop rates of up to 200Hz. Field angles of up-to 2.4 arcmin (4m pupil emulation) can pass through the system. The design is dioptric and permits long cable runs to a compact real-time control system which is on-sky compatible. Therefore experimental validation can be carried out on DRAGON-NG before transferring to an on-sky system, which is a significant risk mitigation.
We have implemented the full AO data-processing pipeline on Graphics Processing Units (GPUs), within the framework of Durham AO Real-time Controller (DARC). The wavefront sensor images are copied from the CPU memory to the GPU memory. The GPU processes the data and the DM commands are copied back to the CPU. For a SCAO system of 80x80 subapertures, the rate achieved on a single GPU is about 700 frames per second (fps). This increases to 1100 fps (1565 fps) if we use two (four) GPUs. Jitter exhibits a distribution with the root-mean-square value of 20 mu s - 30 mu s and a negligible number of outliers. The increase in latency due to the pixel data copying from the CPU to the GPU has been reduced to the minimum by copying the data in parallel to processing them. An alternative solution in which the data would be moved from the camera directly to the GPU, without CPU involvement, could be about 10%-20% faster. We have also implemented the correlation centroiding algorithm, which - when used - reduces the frame rate by about a factor of 2-3.
The Synthetic Aperture Microwave Imaging (SAMI) diagnostic is a Mega Amp Spherical Tokamak (MAST) diagnostic based at Culham Centre for Fusion Energy. The acceleration of the SAMI diagnostic data-processing code by a graphics processing unit is presented, demonstrating acceleration of up to 60 times compared to the original IDL (Interactive Data Language) data-processing code. SAMI will now be capable of intershot processing allowing pseudo-real-time control so that adjustments and optimizations can be made between shots. Additionally, for the first time the analysis of many shots will be possible.
Recent advances in adaptive optics (AO) have led to the implementation of wide field-ofview AO systems. A number of wide-field AO systems are also planned for the forthcoming Extremely Large Telescopes. Such systems have multiple wavefront sensors of different types, and usually multiple deformable mirrors (DMs). Here, we report on our experience integrating cameras and DMs with the real-time control systems of two wide-field AO systems. These are CANARY, which has been operating on-sky since 2010, and DRAGON, which is a laboratory AO real-time demonstrator instrument. We detail the issues and difficulties that arose, along with the solutions we developed. We also provide recommendations for consideration when developing future wide-field AO systems.
Adaptive optics is essential for the successful operation of the future Extremely Large Telescopes (ELTs). At the heart of these AO system lies the real-time control which has become computationally challenging. A majority of the previous efforts has been aimed at reducing the wavefront reconstruction latency by using many-core hardware accelerators such as Xeon Phis and GPUs. These modern hardware solutions offer a large numbers of cores combined with high memory bandwidths but have restrictive input/output (I/O). The lack of efficient I/O capability makes the data handling very inefficient and adds both to the overall latency and jitter. For example a single wavefront sensor for an ELT scale adaptive optics system can produce hundreds of millions of pixels per second that need to be processed. Passing all this data through a CPU and into GPUs or Xeon Phis, even by reducing memory copies by using systems such as GPUDirect, is highly inefficient. The Mellanox TILE series is a novel technology offering a high number of cores and multiple 10 Gbps Ethernet ports. We present results of the TILE-Gx36 as a front-end wavefront sensor processing unit. In doing so we are able to greatly reduce the amount of data needed to be transferred to the wavefront reconstruction hardware. We show that the performance of the Mellanox TILE-GX36 is in-line with typical requirements, in terms of mean calculation time and acceptable jitter, for E-ELT first-light instruments and that the Mellanox TILE series is a serious contender for all E-ELT instruments.
The main goal of Green Flash is to design and build a prototype for a Real-Time Controller (RTC) targeting the European Extremely Large Telescope (E-ELT) Adaptive Optics (AO) instrumentation. The E-ELT is a 39m diameter telescope to see first light in the early 2020s. To build this critical component of the telescope operations, the astronomical community is facing technical challenges, emerging from the combination of high data transfer bandwidth, low latency and high throughput requirements, similar to the identified critical barriers on the road to Exascale. With Green Flash, we will propose technical solutions, assess these enabling technologies through prototyping and assemble a full scale demonstrator to be validated with a simulator and tested on sky. With this R&D program we aim at feeding the E-ELT AO systems preliminary design studies, led by the selected first-light instruments consortia, with technological validations supporting the designs of their RTC modules. Our strategy is based on a strong interaction between academic and industrial partners. Components specifications and system requirements are derived from the AO application. Industrial partners lead the development of enabling technologies aiming at innovative tailored solutions with potential wide application range. The academic partners provide the missing links in the ecosystem, targeting their application with mainstream solutions. This increases both the value and market opportunities of the developed products. A prototype harboring all the features is used to assess the performance. It also provides the proof of concept for a resilient modular solution to equip a large scale European scientific facility, while containing the development cost by providing opportunities for return on investment.
Nigel A. Dipper Durham University E-mail: n.a.dipper@durham.ac.uk A multi-MHz bandwidth two-color interferometer to measure the integral electron density has been developed for the Mega Amp Spherical Tokamak Upgrade (MAST-U). This paper describes the digitization and real time processing part of the system, which is based on field programmable gate arrays (FPGAs) using open source hardware from CERN. The phase shift is measured using heterodyne IQ-downconversion and CORDIC phase rotation enabling detection of integral electron density perturbations above 4MHz. The design is shown to be stable towards 2π-wraps beyond any rate of density change or mirror movement expected in a fusion device. Data is digitized, processed and streamed out in real time with less than 4μs delay making the diagnostic capable of being integrated into real-time Tokamak protection systems and continuousshot experiments.
CuReD (Cumulative Reconstructor with domain Decomposition) and HWR (Hierarchical Wavefront Reconstructor) are novel wavefront reconstruction algorithms for the Shack–Hartmann wavefront sensor, used in the single-conjugate adaptive optics. For a high-order system they are much faster than the traditional matrix–vector-multiplication method. We have developed three methods for mapping the reconstructed phase into the deformable mirror actuator commands and have tested both reconstructors with the CANARY instrument. We find out that the CuReD reconstructor runs stably only if the feedback loop is operated as a leaky integrator, whereas HWR runs stably with the conventional integrator control. Using the CANARY telescope simulator we find that the Strehl ratio (SR) obtained with CuReD is slightly higher than that of the traditional least-squares estimator (LSE). We demonstrate that this is because the CuReD algorithm has a smoothing effect on the output wavefront. The SR of HWR is slightly lower than that of LSE. We have tested both reconstructors extensively on-sky. They perform well and CuReD achieves a similar SR as LSE. We compare the CANARY results with those from a computer simulation and find good agreement between the two.
The next generation of Extremely Large Telescopes (ELTs) for astronomy will rely heavily on the performance of their adaptive optics (AO) systems. Real-time control is at the heart of the critical technologies that will enable telescopes to deliver the best possible science and will require a very significant extrapolation from current AO hardware existing for 4–10 m telescopes. Investigating novel real-time computing architectures and testing their eligibility against anticipated challenges is one of the main priorities of technology development for the ELTs. This paper investigates the suitability of the Intel Xeon Phi, which is a commercial off-the-shelf hardware accelerator. We focus on wavefront reconstruction performance, implementing a straightforward matrix–vector multiplication (MVM) algorithm. We present benchmarking results of the Xeon Phi on a real-time Linux platform, both as a standalone processor and integrated into an existing real-time controller (RTC). Performance of single and multiple Xeon Phis are investigated. We show that this technology has the potential of greatly reducing the mean latency and variations in execution time (jitter) of large AO systems. We present both a detailed performance analysis of the Xeon Phi for a typical E-ELT first-light instrument along with a more general approach that enables us to extend to any AO system size. We show that systematic and detailed performance analysis is an essential part of testing novel real-time control hardware to guarantee optimal science results.
The requirements on the real-time control systems for ELT instruments strongly encourage an investigation of newly emerging hardware and an assessment of its suitability for the job. We have implemented the full AO data processing pipeline on Graphics Processing Units (GPUs), within the framework of Durham AO Real-time Controller (DARC). The pixel data are copied from the CPU memory to the GPU memory. On the GPU, the data are processed and the DM commands are copied back to the CPU. For a system of 80x80 subapertures, the highest rate achieved on a single GPU is 550 frames per second. When running on two or more GPUs, the kernel launching time limits the increase in frame rate. We have also implemented the correlation centroiding algorithm, which - when used - reduces the frame rate by about a factor of two.
DRAGON is a high order, wide field AO test-bench at Durham. A key feature of DRAGON is the ability to be operated at real-time rates, i.e. frame rates of up to 1kHz, with low latency to maintain AO performance. Here, we will present the real-time control architecture for DRAGON, which includes two deformable mirrors, eight wavefront sensors and thousands of Shack-Hartmann sub-apertures. A novel approach has been taken to allow access to the wavefront sensor pixel stream, reducing latency and peak computational load, and this technique can be implemented for other similar wavefront sensor cameras with no hardware costs. We report on experience with an ELT-suitable wavefront sensor camera. DRAGON will form the basis for investigations into hardware acceleration architectures for AO real-time control, and recent work on GPU and many-core systems (including the Xeon Phi) will be reported. Additionally, the modular structure of DRAGON, its remote control capabilities, distribution of AO telemetry data, and the software concepts and architecture will be reported. Techniques used in DRAGON for pixel processing, slope calculation and wavefront reconstruction will be presented. This will include methods to handle changes in CN2 profile and sodium layer profile, both of which can be modelled in DRAGON. DRAGON software simulation techniques linking hardware-in-the-loop computer models to the DRAGON real-time system and control software will also be discussed. This tool allows testing of the DRAGON system without requiring physical hardware and serves as a test-bed for ELT integration and verification techniques.
Nik Looker合作论文数Centre for Advanced Instrumentation at Durham University7