Missing data is a problem commonly seen in most if not all real-world applications. Particularly for water quality monitoring systems, which are commonly plagued by sensor faults or network errors, missing or erroneous data pose a significant challenge in extracting accurate and meaningful insights. In this work, we investigate the problem of missing information in time-series data and propose a new method DISC - Data Imputation with Seasonality and Causality - which uses the concepts of seasonal decomposition and causal discovery to improve contextual accuracy of the imputations for time-series. DISC operates in two stages. First, it builds a causal relational graph representing inter-feature dependencies and uses this graph to impute missing values by adjusting estimates to the nearest-neighbour hourly data points. Second, it learns yearly, monthly, and daily seasonal patterns at an hourly resolution and imputes the remaining gaps. In scenarios where seasonal decomposition fails to fully resolve gaps, causal discovery exploits dependencies among time-series features to generate reference points that enhance the completeness of the seasonal pattern. The proposed method has been evaluated using real-world data from the Murray-Darling Basin and compared with multiple existing machine learning methods. The results validate the effectiveness of DISC, enabling accurate imputation of 14 consecutive days of missing hourly data with an R-squared of 80
The recent success of transformer models for data analysis has resulted in their increased use in varied applications. For some time, transformer models have struggled in Long-Term Time Series Forecasting (LTSF) tasks, where traditional and statistical approaches have been much superior. Recently, a new transformer architecture, iTransformer, has been introduced and has shown state-of-the-art results in various applications. In this paper, we investigate and experiment with the iTransformer model for medium- and long-term forecasting tasks in a real-life application of predicting water temperature in the Murray–Darling Basin, Australia. We compare and evaluate the performance of the iTransformer model to traditional statistical and machine learning approaches. We also investigate different approaches for selecting the historical data for the iTransformer model training and compare the results against a conventional way of selecting the history. Our experimental results show that, in this application, the traditional statistical approach VAR with Seasonal Decomposition (VAR-SD) performs better than the iTransformer model for short-term forecasting range of up to 4 days, but for medium- and long-term forecasting, the iTransformer model performance is superior to both the traditional statistical and machine learning approaches. The results also show that the iTransformer model performs optimally when provided with more recent historical data of approximately the same time length as the forecast horizon.
This paper explores the applicability, efficacy and reliability of various methods for forecasting water temperature of the Murray River in Australia. In particular, Random Forest, XGBoost, ARIMA, ARIMAX, VAR, ARIMA-SD and VAR-SD models have been trained and experimentally evaluated with the real-world data collected from stations along the River. The experimental results show that VAR-SD has produced the best results for short-term (1-6 days) forecasting and that Random Forest is the best for longer-term (7-14 days) forecasting.
This paper presents a simple and intuitive measure of retirement income adequacy: the Pension Multiple which captures and quantifies the level of income for each future retirement year as a multiple of the government-provided social security pension. This Pension Multiple at each future retirement year is then mortality-weighted to produce an average Pension Multiple for the entire retirement. An expected shortfall from this average Pension Multiple is introduced to measure the potential shortfall at each future retirement year from the average Pension Multiple. A single measure of retirement income and potential shortfall is then calculated as the averaged Pension Multiple minus the averaged expected shortfall, which is called the adjusted Pension Multiple. Finally, any residual estate is also included as part of the retirement income adequacy measure. The Pension Multiple and estate residue can be calculated by any forecast model for future retirement income. In this paper, we use a robust stochastic forecasting model to demonstrate the effectiveness of using the Pension Multiple to measure and compare different retirement income strategies.
Despite the importance of drawdown strategies under a defined contribution system with increased longevity risk, little guidance to retired and retiring members has been forthcoming from superannuation funds. This paper provides a do-it-yourself drawdown design for members of superannuation funds along with comparison studies on a range of retirement income strategies under an array of realistic scenarios. A stochastic economic scenario generator is used to simulate the uncertain outcomes of different drawdown strategies during retirement. The impact of annuitisation for mitigating longevity risk under government pension rules and the selection of personalised drawdown and annuitisation strategies for retirement are examined.
The cost of cybersecurity incidents is large and growing. However, conventional methods for measuring loss and choosing mitigation strategies use simplifying assumptions and are often not supported by cyber attack data. In this paper, we present a multivariate model for different, dependent types of attack and the effect of mitigation strategies on those attacks. Utilising collected cyber attack data and assumptions on mitigation approaches, we look at an example of using the model to optimise the choice of mitigations. We find that the optimal choice of mitigations will depend on the goal—to prevent extreme damages or damage on average. Numerical experiments suggest the dependence aspect is important and can alter final risk estimates by as much as 30%. The methodology can be used to quantify the cost of cyber attacks and support decision making on the choice of optimal mitigation strategies.
Recent years have seen a growth in energy research that integrates social and behavioural sciences. A core component of this work involves collecting human data, commonly via surveys and field experiments. But there are often barriers to recruiting large and representative samples of participants, with sampling bias and non-response error posing threats to validity. Identifying cost-effective ways to increase participation in energy research is therefore important for strengthening the rigor, utility and generalisability of studies in this area. To this end, the current study harnesses an experimental design to test pathways for making energy surveys more impactful – specifically by improving response rates and times, lowering sampling bias, and enhancing overall cost-effectiveness. As part of a postal survey on household energy use in Australia, a set of randomised controlled trials were conducted to test the impact of four strategies: incentives, an envelope message, a handwritten sticky note, and a reminder postcard. A 3 x 2 x 2 x 2 factorial design was applied to assess both individual and interactive effects. While material incentives in the form of an upfront token gift and prize draw were ineffective in improving response relative to the control survey, results revealed that a handwritten sticky note expressing upfront thanks for participating – designed to serve as an intrinsically motivating attentional cue – improved both the rate and timeliness of response. Three combinations of strategies yielded significantly higher response rates than the control, but they were more expensive on a ‘dollar cost per response’ basis. Implications for research and practice are discussed.
Abstract The retirement systems in many developed countries have been increasingly moving from defined benefit towards defined contribution system. In defined contribution systems, financial and longevity risks are shifted from pension providers to retirees. In this paper, we use a probabilistic approach to analyse the uncertainty associated with superannuation accumulation and decumulation. We apply an economic scenario generator called the Simulation of Uncertainty for Pension Analysis (SUPA) model to project uncertain future financial and economic variables. This multi-factor stochastic investment model, based on the Monte Carlo method, allows us to obtain the probability distribution of possible outcomes regarding the superannuation accumulation and decumulation phases, such as relevant percentiles. We present two examples to demonstrate the implementation of the SUPA model for the uncertainties during both phases under the current superannuation and Age Pension policy, and test two superannuation policy reforms suggested by the Grattan Institute.
European earthworms have colonised many parts of Australia, although their impact on soil microbial communities remains largely uncharacterised. An experiment was conducted to contrast the responses to Aporrectodea trapezoides introduction between soils from sites with established (Talmo, 64 A. trapezoides m-2) and rare (Glenrock, 0.6 A. trapezoides m-2) A. trapezoides populations. Our hypothesis was that earthworm introduction would lead to similar changes in bacterial communities in both soils. The effects of earthworm introduction (earthworm activity and cadaver decomposition) did not lead to a convergence of bacterial community composition between the two soils. However, in both soils, the Firmicutes decreased in abundance and a common set of bacteria responded positively to earthworms. The increase in the abundance of Flavobacterium, Chitinophagaceae, Rhodocyclaceae and Sphingobacteriales were consistent with previous studies. Evidence for possible soil resistance to earthworms was observed, with lower earthworm survival in Glenrock microcosms coinciding with A. trapezoides rarity in this site, lower soil organic matter and clay content and differences in the diversity and abundance of potential earthworm mutualist bacteria. These results suggest that while the impacts of earthworms vary between different soils, the consistent response of some bacteria may aid in predicting the impacts of earthworms on soil ecosystems.
A wildfire which overran a sensor network site provided an opportunity (a natural experiment) to monitor short-term post-fire impacts (immediate and up to three months post-fire) in remnant eucalypt woodland and managed pasture plots. The magnitude of fire-induced changes in soil properties and soil microbial communities was determined by comparing (1) variation in fire-adapted eucalypt woodland vs. pasture grassland at the burnt site; (2) variation at the burnt woodland-pasture sites with variation at two unburnt woodland-pasture sites in the same locality; and (3) temporal variation pre- and post-fire. In the eucalypt woodland, soil ammonium, pH and ROC content increased post-fire, while in the pasture soil, soil nitrate increased post-fire and became the dominant soluble N pool. However, apart from distinct changes in N pools, the magnitude of change in most soil properties was small when compared to the unburnt sites. At the burnt site, bacterial and fungal community structure showed significant temporal shifts between pre- and post-fire periods which were associated with changes in soil nutrients, especially N pools. In contrast, microbial communities at the unburnt sites showed little temporal change over the same period. Bacterial community composition at the burnt site also changed dramatically post-fire in terms of abundance and diversity, with positive impacts on abundance of phyla such as Actinobacteria, Proteobacteria and Firmicutes. Large and rapid changes in soil bacterial community composition occurred in the fire-adapted woodland plot compared to the pasture soil, which may be a reflection of differences in vegetation composition and fuel loading. Given the rapid yet differential response in contrasting land uses, identification of key soil bacterial groups may be useful in assessing recovery of fire-adapted ecosystems, especially as wildfire frequency is predicted to increase with global climate change.
This paper provides a longitudinal study of withdrawals from account-based pensions from superannuation savings to provide a better understanding of drawdown patterns in retirement. Our analysis of the data indicates that most retirees in their 60s and 70s draw down on their account-based pensions at modest rates, close to the minimum amounts each year. Indeed, if these drawdown rates were to continue, most retirees would die with substantial amounts unspent. These findings are consistent with empirical evidence to date that suggests retirees are inclined to draw down their wealth relatively slowly.
Cellulose accounts for approximately half of photosynthesis-fixed carbon; however, the ecology of its degradation in soil is still relatively poorly understood. The role of actinobacteria in cellulose degradation has not been extensively investigated despite their abundance in soil and known cellulose degradation capability. Here, the diversity and abundance of the actinobacterial glycoside hydrolase family 48 (cellobiohydrolase) gene in soils from three paired pasture-woodland sites were determined by using terminal restriction fragment length polymorphism (T-RFLP) analysis and clone libraries with gene-specific primers. For comparison, the diversity and abundance of general bacteria and fungi were also assessed. Phylogenetic analysis of the nucleotide sequences of 80 clones revealed significant new diversity of actinobacterial GH48 genes, and analysis of translated protein sequences showed that these enzymes are likely to represent functional cellobiohydrolases. The soil C/N ratio was the primary environmental driver of GH48 community compositions across sites and land uses, demonstrating the importance of substrate quality in their ecology. Furthermore, mid-infrared (MIR) spectrometry-predicted humic organic carbon was distinctly more important to GH48 diversity than to total bacterial and fungal diversity. This suggests a link between the actinobacterial GH48 community and soil organic carbon dynamics and highlights the potential importance of actinobacteria in the terrestrial carbon cycle.
Dissolved organic nitrogen (DON) is a significant nitrogen (N) pool in most soils and is considered to be important for N cycling. The present study focused on paired sites of native remnant woodland and managed pasture at three locations in south-eastern Australia. Improved understanding of N cycling is important for assessing the impact of agriculture on soil processes and can guide conservation and restoration soil management strategies to maintain remnant native woodland systems, which currently exist as small pockets of woodland within extensive managed pasture landscapes. Organic and inorganic N pools were quantified, as well as the rates of amino acid and peptide mineralisation in the paired native woodland and managed pasture systems. Soil DON dominated the soil N pool in both land uses, and the proportion of DON to other N pools was greatest at the most N-limited site (up to similar to 70% of extractable N). In both land uses soil ammonium and free amino acid concentrations were similar (similar to 20% of extractable N), and soil nitrate formed the smallest N pool (<similar to 5% of extractable N). Mineralisation of C-14-labelled amino acid and peptide substrates was rapid (<3 h), and more amino acid was respired than peptide in both the native woodland and managed pasture soils. Soil C:N ratio was important in separating site and land use differences, and contrasting relationships between soil physico-chemical properties and organic N uptake rates were identified across sites and land uses. (C) 2015 Elsevier Ltd. All rights reserved.
A new methodology is proposed for the analysis, modeling, and forecasting of data collected from a wireless sensor network. Our approach is considered in the framework of a functional data‐analysis paradigm where observed data is represented in a functional form. To reduce dimensionality, functional principal components analysis is applied to highlight important underlying characteristics and find patterns of variations. The principal scores are modeled with tensor product smooths that allow for smoothing over space and time. The model is then used for simultaneous spatial prediction at unsampled locations and to forecast future observations. We consider soil temperature data from a wireless sensor network of 50 sensor nodes in two different land types (grassland and forest) observed during 60 consecutive days in private property close to Yass, New South Wales, Australia. Copyright © 2015 John Wiley & Sons, Ltd.
Network and multivariate statistical analyses were performed to determine interactions between bacterial and fungal community terminal restriction length polymorphisms as well as soil properties in paired woodland and pasture sites. Canonical correspondence analysis (CCA) revealed that shifts in woodland community composition correlated with soil dissolved organic carbon, while changes in pasture community composition correlated with moisture, nitrogen and phosphorus. Weighted correlation network analysis detected two distinct microbial modules per land use. Bacterial and fungal ribotypes did not group separately, rather all modules comprised of both bacterial and fungal ribotypes. Woodland modules had a similar fungal : bacterial ribotype ratio, while in the pasture, one module was fungal dominated. There was no correspondence between pasture and woodland modules in their ribotype composition. The modules had different relationships to soil variables, and these contrasts were not detected without the use of network analysis. This study demonstrated that fungi and bacteria, components of the soil microbial communities usually treated as separate functional groups as in a CCA approach, were co-correlated and formed distinct associations in these adjacent habitats. Understanding these distinct modular associations may shed more light on their niche space in the soil environment, and allow a more realistic description of soil microbial ecology and function.
Soil heavy metals pollution is an urgent problem worldwide. Understanding the spatial distribution of pollutants is critical for environmental management and decision-making. Children and adults are still routinely exposed to very high levels of heavy metals contaminants in some countries, particularly in regions with a long mining history. In this paper, we analyze lead concentration levels from residential soil samples in the Coeur D’Alene River Basin in the United States. The aim of this paper is to estimate the spatial distribution of the lead concentration levels that may affect exposed humans. Geographic coordinates were compiled for a total of 781 residential addresses and 1,075 mine-related sites (e.g. mine tailings, rock dumps, mine wastes, etc.) surrounding the properties. The lead concentration levels analyzed in the study are in general variable within a residential property and measured levels can differ greatly from one residential address to a nearby address. We consider a unified approach to model the lead concentration levels by means of penalized regression splines and tensor product smooths, using generalized additive models as a building block. We also use this approach to perform a risk assessment spatial analysis to map hot spots for lead based on the action levels defined by the US Environmental Protection Agency.
A problem that frequently arises in environmental surveillance is where to place a set of sensors in order to maximize collected information. In this article we compare four methods for solving this problem: a discrete approach based on the classical k-median location model, a continuous approach based on the minimization of the prediction error variance, an entropy-based algorithm, and simulated annealing. The methods are tested on artificial data and data collected from a network of sensors installed in the Springbrook National Park in Queensland, Australia, for the purpose of tracking the restoration of biodiversity. We present an overview of these methods and a comparison of results.
Four state updating schemes are explored to integrate the observed discharge data into a flood forecasting model. Hourly streamflow discharge measured in the Ovens River catchment, Australia, is assimilated into the Probability Distributed Model (PDM) using the ensemble Kalman filter. The results show that the overall forecast accuracy improves when the discharge observations are integrated, mainly due to better initialisation of the model. Setting error covariance proportional to each state variable gives better results than setting error covariance as a constant value. Updating routing states of PDM affects discharge prediction instantly, while the effect of soil moisture updating results in a lagged response in discharge leading to a poorer update performance. However, during the forecast lead time, updating soil moisture results in slower degradation of the forecast accuracy, which is mainly because the soil moisture store is the only state influencing discharge volume, while the routing storages only describe the flow delay.