The popularity of the CLEAN algorithm in radio interferometric imaging stems from its maturity, speed, and robustness. While many alternatives have been proposed in the literature, none have achieved mainstream adoption by astronomers working with data from interferometric arrays operating in the big data regime. This lack of adoption is largely due to increased computational complexity, absence of mature implementations, and the need for astronomers to tune obscure algorithmic parameters. This work introduces pfb-imaging: a flexible library that implements the scaffolding required to develop and accelerate general radio interferometric imaging algorithms. We demonstrate how the framework can be used to implement a sparsity-based image reconstruction technique known as (unconstrained) SARA in a way that scales with image size rather than data volume and features interpretable algorithmic parameters. The implementation is validated on terabyte-sized data from the MeerKAT telescope, using both a single compute node and Amazon Web Services computing instances.
We report here on studies to determine the accuracy of estimated corrections of Ionospheric Faraday Rotation Measure (IFRM) using observations of the Moon with the Very Large Array (VLA) and MeerKAT telescopes. To estimate the IFRM requires an estimate of the total electron content along the line of sight to the observed sources (the so-called Slant Total Electron Content, or STEC). Estimating the STEC requires an estimate of the global 2D map of Vertical Total Electron Content (VTEC) along with the ray path from the telescope to the source. Traditionally, these global VTEC maps have been utilized along with an assumption that the electrons are in a thin shell at a given altitude to provide an estimate of the IFRM as a function of time. We find that this traditional technique significantly overestimates the IFRM-typically by similar to 0.5-1.1 rad m-2 for the VLA, and similar to-0.3rad m-2 for MeerKAT. Alternatively, the software package ALBUS utilizes raw data from nearby Global Navigation Satellite System stations, to generate a local estimate of the IFRM as a function of time. ALBUS provides considerably better estimates of the IFRM-accurate to similar to 0.1 rad m-2 for both the VLA and MeerKAT, provided the stations utilized have known receiver bias values. A byproduct of our study is the establishment of the intrinsic electric vector position angle of the standard polarized calibrators 3C286 and 3C138 from 500 MHz to 50 GHz, using additional VLA observations of the Moon, Venus, and Mars.
STIMELA2 is a new-generation framework for developing data reduction workflows. It is designed for radio astronomy data but can be adapted for other data processing applications. STIMELA2 aims at the middle ground between ease of development, human readability, and enabling robust, scalable and repeatable workflows. STIMELA2 defines a YAML-based domain specific language (DSL), which represents workflows by linear, concise and intuitive YAML-format recipes. Atomic data reduction tasks (binary executables, Python functions and code, and CASA tasks) are described by YAML-format cab definitions detailing each task's schema (inputs and outputs). The STIMELA2 DSL provides a rich syntax for chaining tasks together, and encourages a high degree of modularity: recipes may be nested into other recipes, and configuration is cleanly separated from recipe logic. Tasks can be executed natively or in isolated environments using containerization technologies such as Apptainer. The container images are open-source and maintained through a companion package called CULT-CARGO. This enables the development of system-agnostic and repeatable workflows. STIMELA2 facilitates the deployment of scalable, distributed workflows by interfacing with the SLURM scheduler and the KUBERNETEs API. The latter allows workflows to be readily deployed in the cloud. Previous papers in this series used STIMELA2 as the underlying technology to run workflows on the AWS cloud. This paper presents an overview of STIMELA2's design, architecture and use in the radio astronomy context.
The physical configuration of new radio interferometers such as MeerKAT, SKA, ngVLA and DSA-2000 informs the development of software in two important areas. Firstly, tractably processing the sheer quantity of data produced by new instruments necessitates subdivision and processing on multiple nodes. Secondly, the sensitivity inherent in modern instruments due to improved engineering practices and greater data quantities necessitates the development of new techniques to capitalize on the enhanced sensitivity of modern interferometers. This produces a critical tension in radio astronomy software development: a fully optimized pipeline is desirable for producing science products in a tractable amount of time, but the design requirements for such a pipeline are unlikely to be understood upfront in the context of artefacts unveiled by greater instrument sensitivity. Therefore, new techniques must continuously be developed to address these artefacts and integrated into a full pipeline. As Knuth reminds us, "Premature optimization is the root of all evil". This necessitates a fundamental trade-off between a trifecta of (1) performant code (2) flexibility and (3) ease-of-development. At one end of the spectrum, rigid design requirements are unlikely to capture the full scope of the problem, while throw-away research code is unsuitable for production use. This work proposes a framework for the development of radio astronomy techniques within the above trifecta. In doing so, we favour flexibility and ease-of-development over performance, but this does not necessarily mean that the software developed within this framework is slow. Practically this translates to using data formats and software from the Open Source Community. For example, by using NuMPy arrays and/or PANDAS dataframes, a plethora of algorithms immediately become available to the scientific developer. Focusing on performance, the breakdown of Moore's Law in the 2010s and the resultant growth of both multi-core and distributed (including cloud) computing, a fundamental shift in the writing of radio astronomy algorithms and the storage of data is required: It is necessary to shard data over multiple processors and compute nodes, and to write algorithms that operate on these shards in parallel. The growth in data volumes compounds this requirement. Given the fundamental shift in compute architecture we believe this is central to the performance of any framework going forward, and is given especial emphasis in this one. This paper describes two Python libraries, DASK-MS and CODEX AFRICANuS which enable the development of distributed High-Performance radio astronomy code with DASK. DASK is a lightweight Python parallelization and distribution framework that seamlessly integrates with the PyDATA ecosystem to address radio astronomy "Big Data" challenges.
Calibration of radio interferometer data ought to be a solved problem; it has been an integral part of data reduction for some time. However, as larger, more sensitive radio interferometers are conceived and built, the calibration problem grows in both size and difficulty. The increasing size can be attributed to the fact that the data volume scales quadratically with the number of antennas in an array. Additionally, new instruments may have up to two orders of magnitude more channels than their predecessors. Simultaneously, increasing sensitivity is making calibration more challenging: low-level RFI and calibration artefacts (in the resulting images) which would previously have been subsumed by the noise may now limit dynamic range and, ultimately, the derived science. It is against this backdrop that we introduce QuartiCal: a new Python package implementing radio interferometric calibration routines. QuartiCal improves upon its predecessor, CubiCal, in terms of both flexibility and performance. Whilst the same mathematical framework - complex optimization using Wirtinger derivatives - is in use, the approach has been refined to support arbitrary length chains of parameterized gain terms. QuartiCal utilizes Dask, a library for parallel computing in Python, to express calibration as an embarrassingly parallel task graph. These task graphs can (with some constraints) be mapped onto a number of different hardware configurations, allowing QuartiCal to scale from running locally on consumer hardware to a distributed, cloud-based cluster. QuartiCal's qualitative behaviour is demonstrated using MeerKAT observations of PSR J2009-2026. These qualitative results are followed by an analysis of QuartiCal's performance in terms of wall time and memory footprint for a number of calibration scenarios and hardware configurations.
Observations of the neutral atomic hydrogen (H i) in the nuclear starburst galaxy NGC 4945 with MeerKAT are presented. We find a large amount of halo gas, previously missed by H i observations, accounting for 6.8 per cent of the total H i mass. This is most likely gas blown into the halo by star formation. Our maps go down to a 3 sigma column density level of 5 x 10(18) cm(-2). We model the H i distribution using tilted-ring fitting techniques and find a warp on the galaxy's approaching and receding sides. The H i in the northern side of the galaxy appears to be suppressed. This may be the result of ionization by the starburst activity in the galaxy, as suggested by a previous study. The origin of the warp is unclear but could be due to past interactions or ram pressure stripping. Broad, asymmetric H i absorption lines extending throughout the H i emission velocity channels are present towards the nuclear region of NGC 4945. Such broad lines suggest the existence of a nuclear ring moving at a high circular velocity. This is supported by the clear rotation patterns in the H i absorption velocity field. The asymmetry of the absorption spectra can be caused by outflows or inflows of gas in the nuclear region of NGC 4945. The continuum map shows small extensions on both sides of the galaxy's major axis that might be signs of outflows resulting from the starburst activity.
We present Tricolour, a package for Radio Frequency Interference mitigation of wideband finely channelized MeerKAT correlation data. The MeerKAT passband is heavily affected by interference from satellite, mobile, aircraft and terrestrial transponders. Coupled with typical data rates in excess of 100 GiB/hr at 208kHz channelization resolution, mitigation poses a significant processing challenge. Our flagger is highly configurable, parallel and optimized, employing Dask and Numba technologies to implement the widely used SumThreshold and MAD interference detection algorithms. We find that typical 208kHz channelized datasets can be processed at rates in excess of 400 GiB/hr for a typical L-band flagging strategy on a modern dual-socket Intel Xeon server.
The next generation of radio telescopes, such as the Square Kilometer Array (SKA), will need to process an incredible amount of data in real-time. In addition, the sensitivity of SKA will require a new generation of calibration and imaging software to exploit its full potential. The wide-field direction-dependent spectral deconvolution framework, called DDFacet, has been successfully used in several existing SKA pathfinders and precursors like MeerKAT and LOFAR. This imager allows a multi-core execution based on facets parallelization and a multinode execution based on observations parallelization. However, because of the amount of data to be processed, the data on a single observation will have to be distributed on several nodes. This paper proposes the first two-level parallelization of DDFacet in the case of a single observation. A multi-core parallelization based on facets and a multi-node parallelization based on frequency distribution grouped in Measurement Sets. We show that this multi-core multi-node parallelization has successfully reduced the total execution time by a factor of 5:7 on a LOFAR dataset.
Superclusters and galaxy clusters offer a wide range of astrophysical science topics with regards to studying the evolution and distribution of galaxies, intra-cluster magnetization mediums, cosmic ray accelerations and large scale diffuse radio sources all in one observation. Recent developments in new radio telescopes and advanced calibration software have completely changed data quality that was never possible with old generation telescopes. Hence, radio observations of superclusters are a very promising avenue to gather rich information of a large-scale structure (LSS) and their formation mechanisms. These newer wide-band and wide field-of-view (FOV) observations require state-of-the-art data analysis procedures, including calibration and imaging, in order to provide deep and high dynamic range (DR) images with which to study the diffuse and faint radio emissions in supercluster environments. Sometimes, strong point sources hamper the radio observations and limit the achievement of a high DR. In this paper, we have shown the DR improvements around strong radio sources in the MeerKAT observation of the Saraswati supercluster by applying newer third-generation calibration (3GC) techniques using CubiCal and killMS software. We have also calculated the statistical parameters to quantify the improvements around strong radio sources. This analysis advocates for the use of new calibration techniques to maximize the scientific returns from new-generation telescopes.
ABSTRACT ESO 149-G003 is a close-by, isolated dwarf irregular galaxy. Previous observations with the ATCA indicated the presence of anomalous neutral hydrogen ($\rm{H{\small I}}$) deviating from the kinematics of a regularly rotating disc. We conducted follow-up observations with the MeerKAT radio telescope during the 16-dish Early Science programme as well as with the MeerLICHT optical telescope. Our more sensitive radio observations confirm the presence of anomalous gas in ESO 149-G003, and further confirm the formerly tentative detection of an extraplanar $\rm{H{\small I}}$ component in the galaxy. Employing a simple tilted-ring model, in which the kinematics is determined with only four parameters but including morphological asymmetries, we reproduce the galaxy’s morphology, which shows a high degree of asymmetry. By comparing our model with the observed $\rm{H{\small I}}$, we find that in our model, we cannot account for a significant (but not dominant) fraction of the gas. From the differences between our model and the observed data cube, we estimate that at least 7–8 per cent of the $\rm{H{\small I}}$ in the galaxy exhibits anomalous kinematics, while we estimate a minimum mass fraction of less than 1 per cent for the morphologically confirmed extraplanar component. We investigate a number of global scaling relations and find that, besides being gas-dominated with a neutral gas-to-stellar mass ratio of 1.7, the galaxy does not show any obvious global peculiarities. Given its isolation, as confirmed by optical observations, we conclude that the galaxy is likely currently acquiring neutral gas. It is either re-accreting gas expelled from the galaxy or accreting pristine intergalactic material.
Xova is a software package that implements baseline-dependent time and channel averaging on Measurement Set data. The uv-samples along a baseline track are aggregated into a bin until a specified decorrelation tolerance is exceeded. The degree of decorrelation in the bin correspondingly determines the amount of channel and timeslot averaging that is suitable for samples in the bin. This necessarily implies that the number of channels and timeslots varies per bin and the output data loses the rectilinear input shape of the input data.
Interferometric coherency measurements scale quadratically with the number of stations in the interferometer. This, combined with high spectro temporal resolution of the data necessitates the use of modern computing strategies such as MapReduce, and cluster computing frameworks to reduce data in tractable amounts of time. Frameworks such as Spark and Dask [1] lean towards a streaming, chunked, functional programming style with minimal shared state. Individual tasks processing chunks of data are flexibly scheduled on multiple cores and nodes. To process the quantities of data produced by contemporary radio telescopes such as MeerKAT, and future telescopes such as the SKA, radio astronomy codes must adapt to these paradigms. In what follows, we describe two Python libraries, dask-ms and codex-africanus, which enable the development of distributed High-Performance Radio Astronomy code.
MeerKATHI is the current development name for a radio-interferometric data reduction pipeline, assembled by an international collaboration. We create a publicly available end-to-end continuum- and line imaging pipeline for MeerKAT and other radio telescopes. We implement advanced techniques that are suitable for producing high-dynamic-range continuum images and spectroscopic data cubes. Using containerization, our pipeline is platform-independent. Furthermore, we are applying a standardized approach for using a number of different of advanced software suites, partly developed within our group. We aim to use distributed computing approaches throughout our pipeline to enable the user to reduce larger data sets like those provided by radio telescopes such as MeerKAT. The pipeline also delivers a set of imaging quality metrics that give the user the opportunity to efficiently assess the data quality.
We present observations and models of the kinematics and the distribution of the neutral hydrogen (HI) in the isolated dwarf irregular galaxy, Wolf-Lundmark-Melotte (WLM). We observed WLM with the Green Bank Telescope (GBT) and as part of the MeerKAT Early Science Programme, where 16 dishes were available. The HI disc of WLM extends out to a major axis diameter of 30 arcmin (8.5 kpc), and a minor axis diameter of 20 arcmin (5.6 kpc) as measured by the GBT. We use the MeerKAT data to model WLM using the TiRiFiC software suite, allowing us to fit different tilted-ring models and select the one that best matches the observation. Our final best-fitting model is a flat disc with a vertical thickness, a constant inclination and dispersion, and a radially-varying surface brightness with harmonic distortions. To simulate bar-like motions, we include second-order harmonic distortions in velocity in the tangential and the vertical directions. We present a model with only circular motions included and a model with non-circular motions. The latter describes the data better. Overall, the models reproduce the global distribution and the kinematics of the gas, except for some faint emission at the 2-sigma level. We model the mass distribution of WLM with a pseudo-isothermal (ISO) and a Navarro-Frenk-White (NFW) dark matter halo models. The NFW and the ISO models fit the derived rotation curves within the formal errors, but with the ISO model giving better reduced chi-square values. The mass distribution in WLM is dominated by dark matter at all radii.
We discuss how the resulting gain solutions can be associated with remaining unmodelled effects, and how deconvolution of extended emission can be improved.