Understanding spatial and temporal patterns of river water quality over a multi-year period is crucial for effective basin management and pollution control. This study applies functional data analysis (FDA) to evaluate monthly water quality index (WQI) data from 16 monitoring stations across the Klang River Basin, Malaysia, covering the period from 2020 to 2023, which spans both pre- and post-pandemic conditions. By treating water quality index (WQI) measurements as smooth functions over time, FDA captures underlying trends and variations that are not readily detected using classical statistical techniques. Functional principal component analysis (FPCA) reveals that the first component accounts for 97% of the total variation, reflecting the dominant pattern in water quality over time, which is characterized by relatively stable upstream conditions and gradual deterioration downstream. The second and third components capture seasonal fluctuations and short-term disturbances, potentially linked to monsoonal cycles and shifts in human activities during the pandemic. Functional clustering based on FPCA scores groups stations according to their temporal behavior, distinguishing upstream areas with stable conditions from downstream areas experiencing greater variability. Spatial interpretation of these clusters offers additional insight into localized pollution sources and environmental stressors. Compared to classical PCA, FDA provides a more detailed, curve-based understanding of time-dependent and location-specific changes in water quality. The result underscore the value of FDA in environmental monitoring, particularly for detecting pre- and post-pandemic shifts, and support its application in guiding adaptive and spatially targeted management strategies for river basins.
The lunar crescent visibility criterion study is one of the most nontrivial discussions that involves astronomy. A lunar crescent visibility criterion is used for calendrical determination, and a different lunar crescent visibility criterion results in different dates of calendar. This paper endeavors to provide an assessment for modern lunar crescent visibility criterion using swarm plot analysis, contradiction rate analysis, and regression analysis, based on 8290 collected data of lunar crescent sighting. The result analysis provides new comparative insight on modern lunar crescent visibility criterion, and suggest a new criterion based on the data of lunar crescent sighting.
Multivariate spatial functional data consists of multiple functions of time-dependent attributes observed at each spatial point. This study focuses on detecting spatial outliers in spatial functional data. Firstly, we develop a new method called Mahalanobis Distance Spatial Outlier (MDSO) to detect functional outliers in the data. The method introduces the multivariate functional Mahalanobis semi-distance and multivariate pairwise functional Mahalanobis semi-distance metrics based on the multivariate functional principal components analysis to calculate the dissimilarity between functions at each spatial point. Via simulation, we show that MDSO performs better than the other competing methods. Secondly, MDSO has been extended to detect spatial functional outliers as well. The functional outliers can now be categorized as global or/and local functional outliers. The appropriate number of neighbors and the cut-off point for the degree of isolation are determined via simulation. Finally, we demonstrate the application of the MDSO on a water quality data set obtained from Sungai Klang basin in Malaysia. The results can be used to support the authority in making better decisions on the management of the river basin or other spatial data with time-independent attributes.
Artificialneural networks (ANNs), modelled after the brain's structure and function, can capture complex nonlinear relationships between predictors and response variables. This study integrates ANNs with the Lee-Carter (LC) framework using a multilayer feed-forward network to forecast the mortality index (k t ), which tracks changes in mortality rates over time within the LC model. This mortality index is essential for forecasting future mortality patterns. Traditionally, the LC model uses an autoregressive integrated moving average (ARIMA) process to predict k t but ARIMA struggles with accurately forecasting future mortality trends. We compared the performance of the multilayer feed-forward network against ARIMA using mortality data from 19 countries, evaluating the models using root mean square error (RMSE) and mean absolute error (MAE). Our findings indicate that the multilayer feed-forward network outperforms ARIMA in forecasting mortality rates for 17 out of 19 countries. Additionally, integrating wavelet analysis and fuzzy logic with ANNs could further enhance forecasting accuracy by effectively managing non-stationary data, uncertainty, and complex patterns.
This study focuses on comparing the performance of the Robust Circular Distance (RCDU*) (simplified version) and A statistics in detecting a single outlier in the Wrapped Normal (WN) samples. Firstly, this study proposes a simplified version of RCDU statistic. Then, the paper generates the cut-off points for both statistics taken from WN samples via a simulation study. This study also evaluates the performance of both statistics using the proportion of a correct outlier detection. As a result, for a small sample size, the performance of RCDU* and A statistics do not have a huge difference. However, for a large sample size of n=250, A statistic performs slightly better than RCDU* statistic. As an illustration of a practical example, both statistics successfully detected one outlier present in the wind direction data at Kota Bharu station.
Abstract Rivers are subject to different sources of pollution. Continuous monitoring of river water quality provides an important basis for the authorities to take appropriate action. Water quality monitoring stations located within the river basin can provide necessary water quality data to establish any changes observed in the river water quality. It is important to highlight lower water quality status at specific monitoring stations so that immediate action can be taken. Similarly, it is an utmost important to ensure water quality at monitoring stations close to water catchment areas always at an acceptable level. This study aims to identify such monitoring stations using descriptive and functional data analysis. The approaches were applied to water quality data collected by the Department of Environment Malaysia at 16 stations in the Klang River basin from January 2013 to December 2016. Specifically, the functional boxplot was applied to identify the monitoring station with outlying properties. We identified many occasions when water quality deteriorated or improved largely due to the increase of COD, BOD and TSS. In addition, three stations close to two main catchment areas and forest reserve showed consistently good water quality. These indicate that the surrounding areas of the stations at the upstream of the rivers are still protected from uncontrolled pollution sources. The study is critical for the authority to understand the overall pattern of water quality data at each station so that action can be planned locally to preserve good river water quality.
The main challenge in the young crescent moon (YCM) observation is the ability to detect the appearance of the YCM, which has varying contrast due to the phenomenon of twilight. The advancement of technology in digital imaging helps the faint and thin image of the YCM to be detected and taken during observation. The techniques used in the observations of the YCM were naked eyes, telescope, and telescope with cameras. A digital imaging technique is also being used in the observations to assist in detecting and recording the image of the YCM more effectively. This paper presents the analysis of the YCM observation data recorded at Telok Kemang Observatory from 2000 to the present. A total of 275 observation sessions were conducted during this study, with 87 positive sightings successfully recorded. The studies found that the smallest elongation and the minimum altitude at sunset of the YCM successfully recorded were 6.81° and 5.40°, respectively. The moon was recorded at an altitude of 3.37°, while the sky is still bright with the sun at an altitude of –2.64° using the digital imaging technique. Based on the records, the YCM which has the minimum criteria of Imkanur Rukyah, i.e., altitude of 2° and elongation of 3° at sunset was never detected or recorded during the 22 years of observations. Therefore, this work suggests the need to change the visibility of Imkanur Rukyah criteria used since 1995 to a more potentially observable criterion. In other aspects, the lengthy observation activities have contributed to the development of a database system for JAKIM that other researchers can access.
In this study, we propose a new method to detect outlying observations in spherical data. The method is based on the k-nearest neighbours distance theory. The proposed method is a good alternative to the existing tests of discordancy for detecting outliers in spherical data. In addition, the new method can be generalized to identify a patch of outliers in the data. We obtain the cut-off points and investigate the performance of the test statistic via simulation. The proposed test performs well in detecting a single and a patch of outliers in spherical data. As an illustration, we apply the method on an eye data set.
The Lee-Carter (LC) model led to the development of many prominent mortality models. This study aims to modify the generalised linear model (GLM) (Poisson, negative binomial, and binomial) framework of the LC model by incorporating factors that affect mortality into the model. The top three factors which affect the mortality for each of the 14 countries studied were selected using the random forest recursive feature elimination (RF-RFE) method which eliminates the least important factors based on the correlation of the predictors with the log-mortality rate. These selected factors were integrated in the form of additional bilinear variates to the GLM models and compared to their original counterparts. The RF-RFE method is effective in selecting the best determinants of mortality by avoiding multicollinearity among predictor variables. The inclusion of the time-factor modulation based on the factors selected improved the model adequacy significantly. Vast improvement was evident in the Poisson and binomial settings. Furthermore, the modified GLM version fits short-base-period data well. This study shows that the inclusion of exogenous determinants of mortality improves the performance of the model significantly.
A spatial outlier refers to the observation whose non-spatial attribute values are significantly different from those of its neighbors. Such observations can also be found in water quality data at monitoring stations within a river network. However, existing spatial outlier detection procedures based on distance measures such as the Euclidean distance between monitoring stations do not take into account the river network topology. In general, water quality levels in lower streams will be affected by the flow from the upper streams. Similarly, the water quality at some tributaries may have little influence on the other tributaries. Hence, a method for identifying spatial outliers in a river network, taking into account the effect of river flow connectivity on the determination of the neighbors of the monitoring stations, is proposed. While the robust Mahalalobis distance is used in both methods, the proposed method uses river distance instead of the Euclidean distance. The performance of the proposed method is shown to be superior using a synthetic river dataset through simulation. For illustration, we apply the proposed method on the water quality data from Sg. Klang Basin in 2016 provided by the Department of Environment, Malaysia. The finding provides a better identification of the water quality in some stations that significantly differ from their neighbouring stations. Such information is useful for the authorities in their planning of the environmental monitoring of water quality in the areas.
This study implements various, maximum overlap, discrete wavelet transform filters to model and forecast the time-dependent mortality index of the Lee-Carter model. The choice of appropriate wavelet filters is essential in effectively capturing the dynamics in a period. This cannot be accomplished by using the ARIMA model alone. In this paper, the ARIMA model is enhanced with the integration of various maximal overlap discrete wavelet transform filters such as the least asymmetric, best-localized, and Coiflet filters. These models are then applied to the mortality data of Australia, England, France, Japan, and USA. The accuracy of the projecting log of death rates of the MODWT-ARIMA model with the aforementioned wavelet filters are assessed using mean absolute error, mean absolute percentage error, and mean absolute scaled error. The MODWT-ARIMA (5,1,0) model with the BL14 filter gives the best fit to the log of death rates data for males, females, and total population, for all five countries studied. Implementing the MODWT leads towards improvement in the performance of the standard framework of the LC model in forecasting mortality rates.
Many astronomers have studied lunar crescent visibility throughout history. Its importance is unquestionable, especially in determining the local Islamic calendar and the dates of important Islamic events. Different criteria have been used to predict the possible visibility of the crescent moon during the sighting process. However, so far, the visibility models used are based on linear statistical theory, whereas the useful variables in this study are in the circular unit. Hence, in this paper, we propose new visibility tests using the circular regression model, which will split the data into three visibility categories; visible to the unaided eye, may need optical aid and not visible. We formulate the procedure to separate the categories using the residuals of the fitted circular regression model. We apply the model on 254 observations collected at Baitul Hilal Teluk Kemang Malaysia, starting from March 2000 to date. We show that the visibility test developed based on elongation of the moon (dependent variable) and altitude of the moon (independent variable) gives the smallest misclassification rate. From the statistical analysis, we propose the elongation of the moon 7.28 degrees, altitude of the moon of 3.33 degrees and arc of vision of 3.74 at sunset as the new crescent visibility criteria. The new criteria have a significant impact on improving the chance of observing the crescent moon and in producing a more accurate Islamic calendar in Malaysia.
Time series of counts occur in many different contexts, the counts being usually of certain events or objects in specified time intervals. In this paper we introduce a model called parameter-driven state-space model to analyse integer-valued time series data. A key property of such model is that the distribution of the observed count data is independent, conditional on the latent process, although the observations are correlated marginally. Our simulation shows that the Monte Carlo Expectation Maximization (MCEM) algorithm and the particle method are useful for the parameter estimation of the proposed model. In the application to Malaysia dengue data, our model fits better when compared with several other models including that of Yang et al. (2015)
We consider the problem of outlier detection method in 2x2 crossover design via Bayesian framework. We study the problem of outlier detection in bivariate data fitted using generalized linear model in Bayesian framework used by Nawama. We adapt their work into a 2x2 crossover design. In Bayesian framework, we assume that the random subject effect and the errors to be generated from normal distributions. However, the outlying subjects come from normal distribution with different variance. Due to the complexity of the resulting joint posterior distribution, we obtain the information on the posterior distribution from samples by using Markov Chain Monte Carlo sampling. We use two real data sets to illustrate the implementation of the method.
Up to now, circular distributions are defined in [0,2 π), except for axial distributions on a semicircle.However, some circular data lie within just half of this range and thus may be better fitted by a half-circular distribution, which we propose and develop in this paper using the inverse stereographic projection technique on a gamma distributed variable.The basic properties of the distribution are derived while its parameters are estimated using the maximum likelihood estimation method.We show the practical value of the distribution by applying it to an eye data set obtained from a glaucoma clinic at the University of Malaya Medical Centre, Malaysia.
Cylindrical data are bivariate data from the combination of circular and linear variables. However, up to now no work has been done on the detection of outlier in cylindrical data. We introduce a definition of outlier for cylindrical data and present a new test of discordancy to detect outlier in this type of data, based on the k-nearest neighbor's distance. Cut-off points of the new test statistic based on the Johnson-Wehrly distribution are calculated and its performance is examined using simulation. A practical example is presented using wind speed and wind direction data obtained from the Malaysian Meteorological Department.
Estimating functions have been used in estimating parameters of many continuous time series models. However, this method has not been applied to models involving count data. In this paper, we use quadratic estimating functions (QEF) to derive estimators for the joint estimation of the conditional mean and variance parameters of count data models, specifically the basic zero-inflated Poisson (ZIP) model, ZIP regression model and integer-valued generalized autoregressive heteroscedastic model with ZIP conditional distribution. Results show that the estimators derived from QEF method, which uses information from combined estimating functions, is more informative than linear estimating functions (LEF) method that only uses information from component estimating functions. Finally, we also fit the real data sets using the ZIP models via QEF, LEF and maximum likelihood methods, and in so doing, demonstrate the superiority of the QEF method in practice.
A cylindrical data set consists of circular and linear variables. We focus on developing an outlier detection procedure for cylindrical regression model proposed by Johnson and Wehrly (1978) based on the k-nearest neighbour approach. The procedure is applied based on the residuals where the distance between two residuals is measured by the Euclidean distance. This procedure can be used to detect single or multiple outliers. Cut-off points of the test statistic are generated and its performance is then evaluated via simulation. For illustration, we apply the test on the wind data set obtained from the Malaysian Meteorological Department.
Gene sequence classification is a well-known problem that impacts several sub-disciplines of Bioinformatics including functional genomics and gene expression data analysis. In gene classification task gene families are frequently formulated using large Generalized Hidden Markov Models (GHMMs) representing a bottleneck for any decoding method and weakening its efficiency. Thus an efficient decoding of such GHMMs remains a key challenge. In this paper, we introduce a new pruned-based strategy for improving the decoding of GHMM using pruning techniques. We focus on viterbi decoding algorithm but the strategy is applicable to GHMM decoding in general. Unlike standard decoding methods, a paradigm shift from screening to-wards recognition is first performed to integrate all considered models into a combined state space. Then the decoding process is limited to the activated states within a beam around the optimal solution to significantly reduce the computational e ort, and thus greatly speeding up the model decoding. Our experiment on Eukaryotic gene demonstrates the e activeness of our approach for speeding up gene classification task.
1 Institute of Mathematical Sciences, Faculty of Science, University of Malaya, 50603 Kuala Lumpur, Malaysia Telekom Research & Development Sdn Bhd., 63000 Cyberjaya, Selangor, Malaysia 2 Institute of Mathematical Sciences, Faculty of Science, University of Malaya, 50603 Kuala Lumpur, Malaysia 3 Department of Mathematics, Faculty of Science, Universiti Putra Malaysia, 43400 UPM, Serdang, Selangor Darul, Ehsan, Malaysia