This study innovates geometric morphometrics by incorporating functional data analysis, the square-root velocity function (SRVF), and arc-length parameterisation for 3D morphometric data, leading to the development of seven new pipelines in addition to the standard geometric morphometrics (GM) approach.. This enables three-dimensional images to be examined from perspectives that do not neglect curvature, through the combined use of arc-length parameterisation, soft-alignment, and elastic-alignment. A simulation study was conducted to demonstrate the general effectiveness of eight pipelines: geometric morphometrics (GM, baseline), arc-GM, functional data morphometrics (FDM), arc-FDM, soft-SRV-FDM, arc-soft-SRV-FDM, elastic-SRV-FDM, and arc-elastic-SRV-FDM. These pipelines were also applied to distinguish dietary categories of kangaroos (omnivores, mixed feeders, browsers, and grazers) using cranial landmarks obtained from 41 extant species. Principal component analysis was conducted, followed by classification analysis using linear discriminant analysis, multinomial regression and support vector machines with a linear kernel. The results highlight the effectiveness of functional data analysis, together with arc-length and SRVF-based approaches, in opening the door to more robust perspectives for analysing three-dimensional morphometrics, while establishing geometric morphometrics as the baseline for comparison.
This study introduces an innovative methodology for mortality forecasting, which integrates signature-based methods within the functional data framework of the Hyndman-Ullah (HU) model. This new approach, termed the Hyndman-Ullah with truncated signatures (HUts) model, aims to enhance the accuracy and robustness of mortality predictions. By utilizing signature regression, the HUts model is able to capture complex, nonlinear dependencies in mortality data which enhances forecasting accuracy across various demographic conditions. The model is applied to mortality data from 12 countries, comparing its forecasting performance against variants of the HU models across multiple forecast horizons. Our findings indicate that overall the HUts model not only provides more precise point forecasts but also shows robustness against data irregularities, such as those observed in countries with historical outliers. The integration of signature-based methods enables the HUts model to capture complex patterns in mortality data, making it a powerful tool for actuaries and demographers. Prediction intervals are also constructed with bootstrapping methods.
We propose a novel extension of the Hyndman-Ullah (HU) model to forecast mortality rates by integrating randomized signatures, referred to as the HU model with randomized signatures (HUrs). Unlike truncated signatures, which grow exponentially with order, randomized signatures, based on the Johnson-Lindenstrauss lemma, are able to approximate higher-order interactions in a computationally feasible way. Using mortality data from four countries, we evaluate the performance of the novel HUrs model compared to two alternative HU model versions. Our empirical results show that the proposed HUrs model performs well, particularly for Bulgarian and Japanese data.
Mortality forecasting is crucial for demographic planning and actuarial studies, especially for projecting population ageing and longevity risk. Classical approaches largely rely on extrapolative methods, such as the Lee-Carter (LC) model, which use mortality rates as the mortality measure. In recent years, compositional data analysis (CoDA), which respects summability and non-negativity constraints, has gained increasing attention for mortality forecasting. While the centred log-ratio (CLR) transformation is commonly used to map compositional data to real space, the α-transformation, a generalisation of log-ratio transformations, offers greater flexibility and adaptability. This study contributes to mortality forecasting by introducing the α-transformation as an alternative to the CLR transformation within a non-functional CoDA model that has not been previously investigated in existing literature. To fairly compare the impact of transformation choices on forecast accuracy, zero values in the data are imputed, although the α-transformation can inherently handle them. Using age-specific life table death counts for males and females in 31 selected European countries/regions from 1983 to 2018, the proposed method demonstrates comparable performance to the CLR transformation in most cases, with improved forecast accuracy in some instances. These findings highlight the potential of the α-transformation for enhancing mortality forecasting within the non-functional CoDA framework.
Amyloid fibrils, characterized by β-sheet-rich protein aggregates, are closely associated with various diseases. Understanding the structural and biochemical changes in amyloid formation requires detailed characterization of their Raman spectroscopic signatures. This study evaluated the application of Raman spectroscopy, utilizing a 532-nm laser excitation source, for differentiating amyloid from normal tissues. Raman spectroscopy effectively identifies protein secondary structures and distinguishes normal tissues from amyloid-containing tissues, offering potential for real-time diagnosis. A total of 13 amyloid tissue samples (heart, kidney, and thyroid) and 9 normal controls were analyzed. Key spectral differences were observed in the amide I (∼1660 cm-1) and amide III (∼1300 cm-1) regions, characteristic of β-sheet structures in amyloid fibrils. Spatially resolved Raman spectra revealed molecular heterogeneity between amide and lipid components in amyloid deposits. Ratiometric analysis further supported this, demonstrating significant differences in the amide-to-lipid ratio (with attributed significant peak intensities at 1660 cm-1 for amide I and 1440 cm-1 for lipids) between amyloid and control tissues. Statistical analysis (Mann-Whitney U test, p = 0.006) confirmed significant differences in amide group intensities between amyloid and control tissues. These findings highlight Raman spectroscopy as a promising tool for real-time identification and characterization of amyloid deposits, with potential clinical applications in diagnosing amyloid-related diseases.
Amyloid diseases are characterized by the accumulation of misfolded protein aggregates in human tissues, pose significant challenges for both diagnosis and treatment. Protein aggregations known as amyloids are linked to several neurodegenerative conditions including Alzheimer's disease, Parkinson's disease, and systemic amyloidosis. The key goal of this research is to employ Small-Angle X-ray Scattering (SAXS) to examine the supramolecular structures of amyloid aggregates in human tissues. We present the structural analysis of amyloid using SAXS, which is employed directly to analyze thin tissue samples without damaging the tissues. This technique provides size and shape information of fibrils, which can be used to generate low-resolution 2D models. The present study investigates the structural changes in amyloid fibril axial d-spacing and scattering intensity in different human tissues, including kidney, heart, thyroid, and others, while also accounting for the presence of triglycerides in these tissues. Tissue structural components were examined at momentum transfer values between q = 0.2 nm- 1 and 1.5 nm- 1. The d-spacing is a critical parameter in SAXS that provides information about the periodic distances between structures within a sample. From the supramolecular SAXS patterns, the axial d-spacing of fibrils in amyloid tissues is prominent and exists within the 3rd to 10th order, compared to that of healthy tissues which do not have notable peak orders. The axial period of fibrils in amyloid tissues is within the scattering vector range 57.40-64.64 nm- 1 while in normal tissues the range is between 60.68 and 61.41 nm-1, which is 3.0 nm- 1 smaller than amyloid-containing tissues. Differences in d-spacing are often correlate with distinct pathological mechanisms or stages of disease progression. The application of SAXS to investigate amyloid structures in human tissues has enormous potential to further knowledge of amyloid disorders. This work will open the path for novel diagnostic instruments and therapeutic strategies meant to reduce the burden of amyloid-related diseases by offering a thorough structural examination of amyloid aggregates.
Increasing the production of the three major food crops (MFCs), maize (Zea mays), rice (Oryza sativa), and wheat (Triticum aestivum), is essential to fulfilling the food demand for the growing human population. Increasing food production may require the integration of machine learning (ML) into plant breeding programs. However, developing ML tools to improve the production of MFCs is a daunting task due to the lack of quality data and the computation resources needed to process this information. Hence, this review discusses the recent applications of ML for improving MFCs production, including plant phenotyping, yield forecasting, and candidate gene prediction. Based on the challenges reported in recent ML experiments for MFCs, this review prescribes solutions to produce scalable ML models. This review provides valuable insights for future studies and promotes collective efforts among researchers implementing ML to enhance MFCs productivity.
In this paper, we explore dimension reduction for functional time series. We propose a generalized dynamic functional principal component analysis (GDFPCA) which does not rely on spectral density estimation and demonstrates strong empirical performance for both stationary and nonstationary functional time series. We define the generalized dynamic functional principal components (GDFPCs) as static factor time series in a functional dynamic factor model and obtain their multivariate representation from a truncation of the functional dynamic factor model. Estimation is based on a least-squares reconstruction criterion and implemented via a two-step procedure for the coefficient vectors of the loading curves under a basis expansion. We establish mean-square consistency of the reconstructed functional time series under weak stationarity. Simulation studies show that GDFPCA performs comparably to dynamic functional principal component analysis (DFPCA) for stationary data, while providing improved reconstruction accuracy in nonstationary settings, where both DFPCA and functional principal component analysis (FPCA) deteriorate. Applications to real datasets support the empirical advantages observed in the simulations.
As is the case for most solid tumours, chemotherapy remains the backbone in the management of metastatic disease. However, the occurrence of chemotherapy resistance is a cause to worry, especially in bladder cancer. Extensive evidence indicates molecular changes in bladder cancer cells to be the underlying cause of chemotherapy resistance, including the reduced expression of farnesyl-diphosphate farnesyltransferase 1 (FDFT1) - a gene involved in cholesterol biosynthesis. This can likely be a hallmark in examining the resistance and sensitivity of chemotherapy drugs. This work performs spectroscopic analysis and metabolite characterization on resistant, sensitive, stable-disease and healthy bladder tissues. Raman spectroscopy has detected peaks at around 1003 cm-1 (squalene), 1178 cm-1 (cholesterol), 1258 cm-1 (cholesteryl ester), 1343 cm-1 (collagen), 1525 cm-1 (carotenoid), 1575 cm-1 (DNA bases) and 1608 cm-1 (cytosine). The peak parameters were examined, and statistical analysis was performed on the peak features, attaining significant differences between the sample groups. Small-angle x-ray scattering (SAXS) measurements observed the triglyceride peak together with 6th, 7th and 8th - order collagen peaks; peak parameters were also determined. Neutron activation analysis (NAA) detected seven trace elements. Carbon (Ca), magnesium (Mg), chlorine (Cl) and sodium (Na) have been found to have the greatest concentration in the sample groups, suggestive of a role as a biomarker for cisplatin resistance studies. Results from the present research are suggested to provide an important insight into understanding the development of drug resistance in bladder cancer, opening up the possibility of novel avenues for treatment through personalised interventions.
This work proposes a functional data analysis approach for morphometrics in classifying three shrew species (S. murinus, C. monticola, and C. malayana) from Peninsular Malaysia. Functional data geometric morphometrics (FDGM) for 2D landmark data is introduced and its performance is compared with classical geometric morphometrics (GM). The FDGM approach converts 2D landmark data into continuous curves, which are then represented as linear combinations of basis functions. The landmark data was obtained from 89 crania of shrew specimens based on three craniodental views (dorsal, jaw, and lateral). Principal component analysis and linear discriminant analysis were applied to both GM and FDGM methods to classify the three shrew species. This study also compared four machine learning approaches (naïve Bayes, support vector machine, random forest, and generalised linear model) using predicted PC scores obtained from both methods (a combination of all three craniodental views and individual views). The analyses favoured FDGM and the dorsal view was the best view for distinguishing the three species.
This work investigates and identifies suitable dimensionality reduction approaches based on variants of principal component analysis (PCA) for various transformations of stock price data. The classical PCA, dynamic principal component analysis (DPCA) and generalised dynamic principal component analysis (GDPCA) were applied to the closing prices, simple returns and log of returns of the top 100 holdings of Standard & Poor's 500 (S&P500) from year 2020 to year 2023. The S&P 500 is a stock market index that tracks the stock performance of 500 large- cap U.S. companies. The performances of the aforementioned variants of PCA on these data for different timeframes were compared. Results showed that GDPCA works best for non- stationary time series data such as the closing prices and DPCA works best for stationary time series data such as the simple returns and the log of returns. The results obtained from the empirical analysis was further supported by simulation studies that follow, hence GDPCA and DPCA could be among the most appropriate dimensionality reduction approaches for non- stationary and stationary time series data respectively.
This study presents the development of multivariate functional Moran's I, along with a novel approach termed multivariate functional areal spatial principal component analysis (mfasPCA), specifically designed for analyzing functional areal data. In addition, we propose a functional permutation-based testing framework that integrates (i) omnibus tests to detect spatial dependence within both positive and negative subspaces, (ii) component wise per-eigen tests that incorporate Holm's method to control the family-wise error rate, and (iii) a sequential rank-wise testing procedure. Through comprehensive simulation studies and an application to empirical data, we demonstrate the efficacy of multivariate functional Moran's I, mfasPCA, and the proposed testing framework in accurately assessing spatial autocorrelation and structural patterns in functional areal data.
Stock market indices are volatile by nature, and sudden shocks are known to affect volatility patterns. The autoregressive conditional heteroskedasticity (ARCH) and generalized ARCH (GARCH) models neglect structural breaks triggered by sudden shocks that may lead to an overestimation of persistence, causing an upward bias in the estimates. Different regime-switching models that have abrupt regime-switching governed by a Markov chain were developed to model volatility in financial time series data. Volatility modelling was also extended to spatially interconnected time series, resulting in spatial variants of ARCH models. This inspired us to propose a Markov switching framework of the spatio-temporal log-ARCH model. In this article, we discuss the Markov-switching extension of the model, the estimation procedure and the smooth inferences of the regimes. The Monte-Carlo simulation studies show that the maximum likelihood estimation method for our proposed model has good finite sample properties. The proposed model was applied to 28 stock indices data that were presumably affected by the 2015-2016 Chinese stock market crash. The results showed that our model is a better fit compared to that of the one-regime counterpart. Furthermore, the smoothed inference of the data indicated the approximate periods where structural breaks occurred. This model can capture structural breaks that simultaneously occur in nearby locations.
Forecasting the financial market has proven to be a challenging task due to high volatility. However, with the growing involvement of computational methods in econometrics, models built with deep learning neural networks have been more accurate in capturing the dynamics of financial market data compared to the commonly used time series models such as the ARIMA and GARCH models. In this study, four deep learning models were applied to eight separate investments, namely stocks (AAPL, TSLA, ROKU, BAC), currency exchange rates (GBP/USD and USD/SEK) and exchange -traded funds (SQQQ and SPXS) to compare their forecasting abilities. The four deep learning models consists of three recurrent neural networks (RNN) which are the vanilla recurrent network (VRNN), long short-term memory (LSTM) and gated recurrent units (GRU), along with the convolutional neural networks (CNN). The models were tuned to be time efficient and evaluated with RMSE and MAPE. Results show that GRU was the overall best model, with exceptions to the LSTM performing better with the exchange traded funds.
In conventional morphometrics, researchers often collect and analyze data using large numbers of morphometric features to study the shape variation among biological organisms. Feature selection is a fundamental tool in machine learning which is used to remove irrelevant and redundant features. Recursive feature elimination (RFE) is a popular feature selection technique that reduces data dimensionality and helps in selecting the subset of attributes based on predictor importance ranking. In this study, we perform RFE on the craniodental measurements of the Rattus rattus data to select the best feature subset for both males and females. We also performed a comparative study based on three machine learning algorithms such as Naïve Bayes, Random Forest, and Artificial Neural Network by using all features and the RFE-selected features to classify the R. rattus sample based on the age groups. Artificial Neural Network has shown to provide the best accuracy among these three predictive classification models.
This work proposes a functional data analysis approach for morphometrics with applications in classifying three shrew species (S. murinus, C. monticola and C. malayana) based on the images. The discrete landmark data of craniodental views (dorsal, jaw and lateral) are converted into continuous curves where the curves are represented as linear combinations of basis functions. A comparative study based on four machine learning algorithms such as naive Bayes, support vector machine, random forest, and generalized linear models was conducted on the predicted principal component scores obtained from the FDA approach and classical approach (combination of all three craniodental views and individual views). The FDA approach produced better results in separating the three clusters of shrew species compared to the classical method and the dorsal view gave the best representation in classifying the three shrew species. Overall, based on the FDA approach, GLM of the predicted PCA scores was the most accurate (95.4% accuracy) among the four classification models.
This work focuses on functional data presenting spatial dependence. The spatial autocorrelation of stock exchange returns for 71 stock exchanges from 69 countries was investigated using the functional Moran’s I statistic, classical principal component analysis (PCA) and functional areal spatial principal component analysis (FASPCA). This work focuses on the period where the 2015–2016 global market sell-off occurred and proved the existence of spatial autocorrelation among the stock exchanges studied. The stock exchange return data were converted into functional data before performing the classical PCA and FASPCA. Results from the Monte Carlo test of the functional Moran’s I statistics show that the 2015–2016 global market sell-off had a great impact on the spatial autocorrelation of stock exchanges. Principal components from FASPCA show positive spatial autocorrelation in the stock exchanges. Regional clusters were formed before, after and during the 2015–2016 global market sell-off period. This work explored the existence of positive spatial autocorrelation in global stock exchanges and showed that FASPCA is a useful tool in exploring spatial dependency in complex spatial data.
Trace and minor elements play crucial roles in a variety of biological processes, including amyloid fibrils formation. Mechanisms include activation or inhibition of enzymatic reactions, competition between elements and metal proteins for binding positions, also changes to the permeability of cellular membranes. These may influence carcinogenic processes, with trace and minor element concentrations in normal and amyloid tissues potentially aiding in cancer diagnosis and etiology. With the analytical capability of the spectroscopic technique X-ray fluorescence (XRF), this can be used to detect and quantify the presence of elements in amyloid characterization, two of the trace elements known to be associated with amyloid fibrils. In present work, involving samples from a total of 22 subjects, samples of normal and amyloid-containing tissues of heart, kidney, thyroid, and other tissue organs were obtained, analyzed via energy-dispersive X-ray fluorescence (EDXRF). The elemental distribution of potassium (K), calcium (Ca), arsenic (As), and iron (Fe) was examined in both normal and amyloidogenic tissues using perpetual thin slices. In amyloidogenic tissues the levels of K, Ca, and Fe were found to be less than in corresponding normal tissues. Moreover, the presence of As was only observed in amyloidogenic samples; in a few cases in which there was an absence of As, amyloid samples were found to contain Fe. Analysis of arsenic in amyloid plaques has previously been difficult, often producing contradictory results. Using the present EDXRF facility we could distinguish between amyloidogenic and normal samples, with potential correlations in respect of the presence or concentration of specific elements.
Abstract This research introduces a new method for analysing shape variation for 3D landmark coordinate data, called functional data geometric morphometrics (FDGM). FDGM uses functional data analysis (FDA) to treat landmark coordinates as continuous curves or functions. This allows for a more exhaustive description and analysis of shape variation compared to geometric morphometrics (GM), which treats landmark coordinates as discrete points. A simulation study was conducted to demonstrate the general effectiveness of FDGM compared to the GM. Principal component analysis (PCA) and linear discriminant analysis (LDA) were applied to both the landmark coordinates and the functional form of the landmark coordinates. The analyses favoured FDGM. The reconstruction error for FDGM was smaller when smoothed data was considered in generating the data. FDGM and GM were then applied to distinguish dietary categories of kangaroos (omnivores, mixed feeders, browser, and grazer) using landmarks obtained from crania of 41 kangaroo extant species. The results demonstrate that FDGM is a powerful method for analysing shape variation in 3D landmark coordinate data. FDGM can substantially enhance the domain of morphometrics, providing a valuable resource for driving future progress within this realm.