We propose a measure, the joint differential entropy of eigencolors, for determining the spatial complexity of exoplanets using only spatially unresolved light-curve data. The measure can be used to search for habitable planets, based on the premise of a potential association between life and exoplanet complexity. We present an analysis using disk-integrated light curves from Earth, developed in previous studies, as a proxy for exoplanet data. We show that this quantity is distinct from previous measures of exoplanet complexity due to its sensitivity to spatial information that is masked by features with large mutual information between wavelengths, such as cloud cover. The measure has a natural upper limit and appears to avoid a strong bias toward specific planetary features. This makes it a novel and generalizable method, which, when combined with other methods, can broaden the available indicators of habitability.
We present a novel natural language processing (NLP) approach to deriving plain English descriptors for science cases otherwise restricted by obfuscating technical terminology. We address the limitations of common radio galaxy morphology classifications by applying this approach. We experimentally derive a set of semantic tags for the Radio Galaxy Zoo EMU (Evolutionary Map of the Universe) project and the wider astronomical community. We collect 8,486 plain English annotations of radio galaxy morphology, from which we derive a taxonomy of tags. The tags are plain English. The result is an extensible framework which is more flexible, more easily communicated, and more sensitive to rare feature combinations which are indescribable using the current framework of radio astronomy classifications.
Large area astronomical surveys will almost certainly contain new objects of a type that have never been seen before. The detection of 'unknown unknowns' by an algorithm is a difficult problem to solve, as unusual things are often easier for a human to spot than a machine. We use the concept of apparent complexity, previously applied to detect multi-component radio sources, to scan the radio continuum Evolutionary Map of the Universe (EMU) Pilot Survey data for complex and interesting objects in a fully automated and blind manner. Here we describe how the complexity is defined and measured, how we applied it to the Pilot Survey data, and how we calibrated the completeness and purity of these interesting objects using a crowd-sourced 'zoo'. The results are also compared to unexpected and unusual sources already detected in the EMU Pilot Survey, including Odd Radio Circles, that were found by human inspection.
Prediction of Total Cloud Cover (TCDC) from numerical weather simulation models, such as Global Forecast System (GFS), can aid renewable energy engineers in monitoring and forecasting solar photovoltaic power generation. A major challenge is the systematic bias in TCDC simulations induced by the errors in the numerical model parameterization stages. Correction of GFS-derived cloud forecasts at multiple time steps can improve energy forecasts in electricity grids to bring better grid stability or certainty in the supply of solar energy. We propose a new kernel ridge regression (KRR) model to reduce bias in TCDC simulations for medium-term prediction at the inter-daily, e.g., 2–8 day-ahead predicted TCDC values. The proposed KRR model is evaluated against multivariate recursive nesting bias correction (MRNBC), a conventional approach and eight machine learning (ML) methods. In terms of the mean absolute error (MAE), the proposed KRR model outperforms MRNBC and ML models at 2–8 day ahead forecasts, with MAE ≈ 20–27%. A notable reduction in the simulated cloud cover mean bias error of 20–50% is achieved against the MRNBC and reference accuracy values generated using proxy-observed and non-corrected GFS-predicted TCDC in the model’s testing phase. The study ascertains that the proposed KRR model can be explored further to operationalize its capabilities, reduce uncertainties in weather simulation models, and its possible consideration for practical use in improving solar monitoring and forecasting systems that utilize cloud cover simulations from numerical weather predictions.
ABSTRACT The Evolutionary Map of the Universe (EMU) large-area radio continuum survey will detect tens of millions of radio galaxies, giving an opportunity for the detection of previously unknown classes of objects. To maximize the scientific value and make new discoveries, the analysis of these data will need to go beyond simple visual inspection. We propose the coarse-grained complexity, a simple scalar quantity relating to the minimum description length of an image that can be used to identify unusual structures. The complexity can be computed without reference to the broader sample or existing catalogue data, making the computation efficient on new surveys at very large scales (such as the full EMU survey). We apply our coarse-grained complexity measure to data from the EMU Pilot Survey to detect and confirm anomalous objects in this data set and produce an anomaly catalogue. Rather than work with existing catalogue data using a specific source detection algorithm, we perform a blind scan of the area, computing the complexity using a sliding square aperture. The effectiveness of the complexity measure for identifying anomalous objects is evaluated using crowd-sourced labels generated via the Zooniverse.org platform. We find that the complexity scan identifies unusual sources, such as odd radio circles, by partitioning on complexity. We achieve partitions where 5 per cent of the data is estimated to be 86 per cent complete, and 0.5 per cent is estimated to be 94 per cent pure, with respect to anomalies and use this to produce an anomaly catalogue.
We define deriving semantic class targets as a novel multi-modal task. By doing so, we aim to improve classification schemes in the physical sciences which can be severely abstracted and obfuscating. We address this task for upcoming radio astronomy surveys and present the derived semantic radio galaxy morphology class targets.
The volume of data that will be produced by the next generation of astrophysical instruments represents a significant opportunity for making unplanned and unexpected discoveries. Conversely, finding unexpected objects or phenomena within such large volumes of data presents a challenge that may best be solved using computational and statistical approaches. We present the application of a coarse-grained complexity measure for identifying interesting observations in large astronomical data sets. This measure, which has been termed apparent complexity, has been shown to model human intuition and perceptions of complexity. Apparent complexity is computationally efficient to derive and can be used to segment and identify interesting observations in very large data sets based on their morphological complexity. We show, using data from the Australia Telescope Large Area Survey, that apparent complexity can be combined with clustering methods to provide an automated process for distinguishing between images of galaxies which have been classified as having simple and complex morphologies. The approach generalizes well when applied to new data after being calibrated on a smaller data set, where it performs better than tested classification methods using pixel data. This generalizability positions apparent complexity as a suitable machine-learning feature for identifying complex observations with unanticipated features.