
An important tool in the Genomic approach is the widely used Microarray interventions to effectively and accurately predict survival time and risk analysis of patients. This study is aimed to present and examine the risk impact of five cancer gene sequences with different high-dimensional microarray data sets, through survival analysis. GSE10300; GSE14333; GSE16446; GSE17618 and GSE20685 data sets with 16,183; 54,712; 54,739; 54,726 and 54,739 gene counts from sample sizes 44, 226, 107, 44 and 327 respectively were selected from the NCBI database to identify their risk impact. Then the classical Cox Proportional Hazard (CoxPH), Bayesian Adaptive Spline Surface (BASS) and Bayesian Additive Regression Tree (BART) models were employed to analyze each of the five data sets. Each model’s prediction performance was measured using the Root Mean Squared Error (RSME). The fol-lowing, R packages ‘survive’ for CoxPH; ‘BASS’ for BASS and ‘bartMachine’ for BART were utilized in obtaining the results. Finally, influential variable plots were shown to indicate the most important main effects and interactions for the five microarray datasets. Varying median hazard values and traces of yellow boxes across the diagonals of the heatmap diagrams indicate multi-collinearity between genes, implying misleading results by the CoxPH model, while leveraging superiority of BART model over the BASS and CoxPH models. The RSME plots successfully identify some influential genes among thou-sands of genes across the five real-life microarray datasets analyzed: CoxPH model shows a moderate level of predictive accuracy while the BASS model exhibits both strengths and weaknesses depending on the number of predictors used, significantly higher than CoxPH. The BART model, on the other hand, with the lowest RSME, consistently demonstrates its superior, stable and consistent predictive performance over CoxPH and BASS models. This study gives a more distinct and nuanced understanding of genes interactions and influence, as well as provide improved prognostic tools and personalized treatment strategies for cancer patients’ risks and survival.
Agriculture is a vital sector in Sri Lanka, yet it has been increasingly challenged by environmental disturbances and policy shifts. This study analyzes the impact of two significant disruptions—the 2017 drought and the 2021 fertilizer ban—on paddy yield trends. To address the complexities of these influences, Bayesian spline regression was employed to model non-linear and seasonal yield patterns, and interrupted time series regression was used to capture abrupt changes in response to external shocks. Uncertainty and the causal effects of these interventions were quantified through posterior distributions, estimated using the brms package and the NUTS algorithm. The results offer nuanced insights into how environmental and policy factors shape agricultural productivity, with implications for researchers and policymakers concerned with sustainable agriculture and climate adaptation. Overall, this study underscores the value of Bayesian modeling in assessing multifaceted, real-world issues that span agriculture, environment, and policy.
This study investigates the impact of major national crises on life expectancy trends in Sri Lanka from 2000 to 2023 using a Bayesian generalized additive modeling (GAM) framework combined with interrupted time series analysis. The model incorporates key socio-economic and environmental covariates—including infant mortality and GDP per capita—along with event indicators to capture the immediate and post-event effects of significant crises: the 2004 Indian Ocean tsunami, the prolonged civil conflict culminating in 2009, and the COVID-19 pandemic. Model evaluation based on the Leave-One-Out Information Criterion (LOOIC) identified a spline dimension of 8 as optimal for capturing nonlinear temporal trends. Results indicate significant immediate negative impacts on life expectancy due to the 2004 tsunami and the 2009 war. The nonlinear temporal trend, captured by the spline term, showed a strong positive effect indicating overall improvements over time. Estimated effects for infant mortality and GDP per capita were inconclusive, with wide credible intervals overlapping zero, suggesting minimal or uncertain direct associations after accounting for crises and temporal trends. Posterior predictive checks confirmed the model’s strong predictive accuracy. Forecasting for 2024 was performed using a second-order Taylor series expansion on posterior fitted draws, providing a robust and computationally efficient approach to project life expectancy beyond observed data while accounting for nonlinear trends and uncertainty. The forecast indicates a continued upward trend, though with increasing uncertainty. These findings highlight the resilience of the Sri Lankan health system and underscore the importance of targeted interventions during crises to sustain population health improvements. The proposed modeling framework provides policymakers a valuable tool for ongoing health monitoring, crisis impact assessment, and strategic planning in Sri Lanka and similar settings.
In this paper, we have given four characterization results of two families of continuous probability distribution. The conditional expectation of function of order statistics is used to establish the required results when the conditioned order statistics may not be adjacent one. Further, its important deduction is also discussed.
Dengue fever imposes a substantial economic burden on healthcare systems, particularly in low- and middle-income countries. This study aimed to assess the financial impact of dengue fever in a private hospital setting in Sri Lanka and to evaluate the predictive value of on-admission atypical lymphocyte count (ALC) for disease severity, length of hospital stay, and total hospital costs. A retrospective analysis was conducted on 2,185 confirmed dengue patients admitted to Nawaloka Hospital, Colombo, between January 2017 and August 2023. Data on patient demographics, disease severity (per WHO 2009 classification), financial records, and ALC values from the Sysmex XS500i analyzer were extracted. The median total hospital bill was LKR 127,600, with room charges and administrative costs contributing the most. Severe dengue, prolonged hospitalization (>5 days), and insurance coverage were significantly associated with higher hospital bills (p<0.001). Patients with an ALC >0.5×103/μL on day 3 of fever had significantly longer hospital stays and higher costs compared to those with lower ALC levels. ROC analysis showed that elevated ALC predicted hospital stays >5 days with 94.4% sensitivity and 78.4% specificity, and costs >LKR 132,000 with 95.8% sensitivity and 72.4% specificity. These findings highlight ALC as a practical, early biomarker to stratify dengue patients by risk and optimize resource al-location. The study underscores the need for early clinical decision-making o minimize economic strain on patients, especially those lacking insurance coverage. Further prospective research is recommended to validate ALC as a cost-saving predictive tool in dengue management.
In this paper, we propose a modified regression exponential estimator of population mean of the sensitive variable for additive model of Randomized Response Technique (RRT) given by Pollack and Bek (1976). The proposed estimator utilizes a sufficiently correlated non-sensitive auxiliary variable to improve estimation accuracy. We derive the expression for bias and mean square error (MSE) of the estimator up to the first order of approximation. Furthermore, an empirical and theoretical study is conducted using artificially generated populations for comparing efficiency of the pro-posed estimator with existing estimators. The results demonstrate that the present relative efficiency of the proposed estimator generally is comparison of the other existing estimators for different values of correlation coefficient.
Accurate estimation of rare events in clustered populations is paramount for effective public health planning, yet traditional estimators often assume population homogeneity, leading to inefficiencies and potential bias. This study employs a case study approach using a real-world public health dataset to compare two estimators for the proportion of a rare event (incomplete Hepatitis B vaccination) within an Adaptive Cluster Sampling (ACS) frame-work: a classical ACS estimator and a heterogeneous estimator derived from a Bayesian hierarchical model. The hierarchical structure accounts for clustering at the College-Department level (networks, representing 279 distinct net-work units identified through the adaptive sampling process, including both expanded clusters and singleton networks) and the College-Year level (sub-networks, representing 45 nested groupings within the 380-student population). By treating a complete dataset as a known (population), we rigorously assess estimator performance in terms of bias, variance, and Mean Squared Error (MSE). Our findings demonstrate that the Bayesian hierarchical estimator (estimate: 0.182, variance: 0.00038, MSE: 0.00038) consistently yields estimates with negligible bias, lower variance, and a substantially lower MSE compared to the classical ACS estimator (estimate: 0.172, variance: 0.00050, MSE: 0.00059). This represents a 1.31-fold reduction in variance and a 1.55-fold reduction in MSE for the Bayesian approach. Posterior predictive checks further confirm the good fit of the Bayesian model to the observed data. This underscores the critical importance of explicitly accounting for population heterogeneity and employing robust model-based inference in adaptive sampling designs to generate reliable estimates for informed public health policy.
In this study, we introduced a novel generalized class of ratio-type estimators for median estimation by employing quantile regression within the framework of neutrosophic statistics is designed to enhance the accuracy and reliability of estimates in the presence of data uncertainty, providing a robust alternative to traditional point estimators. The proposed methodology yields interval-based estimates for the population median, capturing a range of possible values rather than a point estimate, which allows modeling uncertainty and partial truth inherent in complex or imprecise data. This is achieved within the neutrosophic statistical framework, a generalization of classical statistics that explicitly accounts for indeterminacy and inconsistent information, making it well-suited for handling ambiguous data in interdisciplinary applications. We validate the performance of our estimators using real-life stock price data from Samsung Electronics Co. Ltd. (SMSN.IL), sourced from Yahoo Finance (2022), and further substantiate the results through a comprehensive simulation study. Comparisons between traditional and proposed estimators, based on mean squared error, percentage decrease in mean squared error, and a simulation study, demonstrate the superior precision and reliability of our approach.
The application of calibration in domain estimation with subsampling the non-respondents is the major emphasis of this work. The detrimental impact that result from a researcher’s incapacity to gather data from every elemen-tary unit in a certain demographic or from respondents’ partial or complete unwillingness to supply the necessary information is still a problem in sample surveys. In view of the mentioned problem, this work becomes imperative as it utilizes the concept of calibration in double sampling. In order to increase the accuracy of such estimates by subsampling the non-respondents, it becomes important to close the gap between the loss of survey data and establishing a reliable estimate in the research areas. This formulation is subject to two conditions, when the auxiliary variable is free from non-response (condition A) and when the auxiliary variable is not free from non-response (condition B). And it has been demonstrated that the suggested estimators are unbiased. Two populations from actual data were examined in terms of variance/MSE and percentage relative efficiency (PRE) as part of the empirical study. Ad-ditionally, two non-response scenarios-one with uniform rates and the other with non-uniform rate were taken into account. The findings demonstrated that although the two conditions yield the same population mean estimates, condition A is more efficient than condition B. The cases, on the other hand, produce different estimates of population mean and variance/MSE. Compared to the existing estimators used in this work, the proposed estimator has proven to be more effective in providing accurate estimates of the population mean in the domains. This is evident from the efficiency comparisons and empirical investigations of the proposed estimator and the existing estimators that the proposed estimator are capable of producing efficient estimates in domains of study when subsampling the nonrespondents.
Health outcomes research assesses and evaluates the result of health interventions, care, and therapy. Health outcomes research studies often use longitudinal data or repeated measures of health outcomes from interventions as dynamic measures. However, the traditional regression methods of analyzing such data are inherently limited in scope and flexibility. Latent growth models provide the needed flexibility and robustness for longitudinal data analysis. To illustrate the rationale and methodology of latent growth models in the context of outcomes research, the author used simulated hypothetical data of 357 hypertensive patients with controlled blood pressure attending a clinic with blood pressure measurements collected over four time points or waves (T1, T2, T3, and T4). The growth model was developed using structural equation modeling with Analysis of Moment Structures (AMOS) software. The study assumed that blood pressure control followed a linear trend. The results showed that the simulated model provided an adequate fit to the data. At the intraindividual level, the mean blood pressure measure (intercept) at the starting point was significant. The rate of change (slope or growth parameter) was positive and significant, indicating improvement of blood pressure measures from one time point to another. The correlation between the intercept and slope parameters was not substantial, implying that the initial measure did not necessarily change over time (consistency). To explain possible interindividual differences based on assumed significant variances of the error variances of intercept and slope parameters, the effect of a time-invariant variable-sex on the conditional model was explored. The use of latent growth curve modeling in health outcomes research can enrich the scope and quality of information obtained from longitudinal outcomes data. Reporting format, interpretation, and further implications for outcomes research are provided in the paper.
The proposed research incorporates the utilization of a heavy-tailed skewe–d distribution referred to as the inverse Weibull as a link function in the context of a binary classification model. This selection is motivated by the need to ad-dress the existence of rare or extreme events in random processes. The study introduces a model that relies on the Inverse Weibull (TYPE II) distribution, and the estimation of model parameters is accomplished through the appli-cation of maximum likelihood measures. When the results are compared to those derived from other link functions such as TYPE I (Complementary log) and TYPE III (Weibull) based on extreme value distributions using simulation data, it becomes apparent that the Inverse Weibull (TYPE II) model exhibits exceptional performance. This performance assessment takes into account several criteria, including the Akaike information criterion, the Bayesian in-formation criterion, the area under the curve, and the Brier scores. In conclu-sion, the study establishes that the proposed model demonstrates considerable robustness in its performance, rendering it a viable choice for the modeling of binary classification problems.
Forecasting cryptocurrency prices is challenging due to their highly vola–tile and unpredictable nature. This study presents a high-frequency forecasting model that combines changepoint analysis with clustering-based machine lear–ning techniques. The pruned exact linear time (PELT) method was applied to intraday historical data from Bitcoin (BTC), Ethereum (ETH), and Cardano (ADA) for a one-year period, from August 2022 to August 2023, to detect volatility shifts, dividing the data into distinct regions. These regions were then grouped using the K-means algorithm based on characteristics like price and volume variance. Support Vector Regression (SVR), Long Short-Term Memory (LSTM), and Random Forest (RF) models were trained on both clustered and non-clustered data to predict cryptocurrency closing prices. The results showed that using clustered original data improved forecasting accuracy in most cases. For ADA, SVR’s Mean Absolute Error (MAE) dropped from 0.0018 to 0.0003, and for RF, MAE improved from 0.0010 to 0.0004. Similar improvements were observed for BTC and ETH. These results show that clustering volatility regions enables machine learning models to more accurately capture price dynamics, resulting in better and more reliable forecasts. This study demonstrates the effectiveness of integrating changepointbased clustering with machine learning models to improve short-term cryptocurrency price forecasts.
Multiple Hypothesis Testing presents challenges due to increased false discoveries when conducting statistical tests simultaneously. Despite the pro-posal of novel correction techniques, the procedure of selecting the most suit-able method has remained a black box. The trade-off that arises from con-trolling false positives and negatives through different correction techniques underlines the need for a cohesive framework. This study addresses the above challenges with a special focus on gene expression data and evaluates six widely used Multiple Hypothesis Testing (MHT) methods, namely, Bonferroni, Holm, sequential goodness of fit (SGoF), Benjamini-Hochberg, Benjamini-Yekutieli, and Storey’s Q-values, across different scenarios to compare two in-dependent groups. Our results show that Storey’s Q-value performs well with large effect sizes, whereas SGoF excels in low-effect scenarios. However, the Bonferroni and Holm methods offer high precision owing to the strict control of false positives. Recognizing the limitations of relying on a single method, we introduce a novel Weighted Hybrid Method (WHM), a decision-support framework that allows users to navigate between approaches rather than serv-ing as a new statistical test. An innovative Significant Index Plot (SIP) is un-veiled to assist in the detection of significant hypotheses across different meth-ods. The framework was tested on four genomic datasets: gene expression in multiple sclerosis (GSE21942), myelodysplastic syndrome (GSE61853), alcohol-related gene expression (GSE52553), and age-related corneal tran-scriptomes (GSE58315), extending the usage to enable independent hypoth-esis weighting. A novel Python library, MultiDST, and a web interface were developed to enable researchers to apply the framework efficiently, improving the transparency of their findings.
Fresh-cut vegetables (FCVs) offer convenience and nutritional benefits, meeting the demands of today’s fast-paced lifestyle. While the Sri Lankan FCV market is still emerging, many developed countries have already embraced FCVs. This study aimed to identify consumer profiles in urban communities in Sri Lanka based on their preferences in FCVs and purchasing concerns. A cross-sectional survey was conducted using both face-to-face interviews and online questionnaires, resulting in 1722 responses. Descriptive and inferential statistics were employed, including correspondence analysis, Kruskal–Wallis test, factor analysis, and cluster analysis. The findings revealed that convenience, freshness, and affordability are key factors motivating the choice of FCVs over whole vegetables. Consumers showed a strong preference for fruit-type FCVs, followed by root vegetables, while bulbs, tubers, and chillies were among the least preferred. Among labeling information, expiry date and preservation method were prioritized, while nutritional information was less emphasized. Correspondence analysis indicated that natural disinfectants (e.g., salt, vinegar, and baking soda) are more commonly associated with plastic and rigifoam trays. Factor analysis reduced the 12 purchasing concerns into 03 key components as quality, safety, and presentation, explaining approximately 62% of the total variance. Based on standardized factor scores, consumer segmentation was performed through hierarchical clustering fol-lowed by K-means clustering. The analysis revealed three distinct consumer segments: 1) less concerned about all three factors (the smallest cluster); 2) moderately concerned about all three factors and quality-driven consumers; and 3) highly concerned about all three factors and safety-driven consumers. These insights offer practical guidance for advancing the commercial viability of FCVs in Sri Lanka. Future research will expand to include non-urban populations to improve generalizability.
In time series analysis, external shocks are essential for accurately capturing the effects of sudden, non-recurring events, enabling more precise estimation and forecasting. This study examines Sri Lankan economic variables including monthly year-on-year inflation rates, exchange rates, and interest rates using monthly data from 2003 to 2023 obtained from the International Monetary Fund (IMF) database and the Central Bank of Sri Lanka. Dummy variables for natural disasters, health crises, and political events were included to account for external shocks. Time series plots of inflation, exchange rates, and interest rates were generated to identify trends, seasonal patterns, and anomalies. Cointegration of the residuals of the fitted models was tested, revealing the presence of cointegration. Three Vector Error Correction (VEC) models were estimated, with one model selected based on information criteria and model diagnostic techniques. Model diagnostics ensured the reliability of the VEC models, and the VEC model with log transformed time series variables with dummy variables was chosen due to its lowest Mean Squared Error (MSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE). Diagnostic tests, including the Portmanteau test for serial correlation and stability checks, were conducted. The Portmanteau test revealed significant autocorrelation in the residuals, indicating that the VEC model may not fully capture the data’s dynamics and might need modifications. However, stability analysis confirmed that all eigenvalues of the selected VEC model are within the unit circle, ensuring reliable forecasts. The selected VEC model provides valuable insights into Sri Lanka’s economic variables. Nonetheless, the findings suggest that further adjustments may be needed to address residual autocorrelation.
Inventory management is a crucial research area focusing on stock optimization, cost reduction, and supply chain efficiency, with particular attention given to optimizing practices and identifying key factors. If inventory management fails to maintain a sustainable level, it directly hampers the company. To address this problem, we propose a new methodology that combines Bayesian regression analysis with optimization techniques. The historical sales data used in this study was obtained from an online database platform called “Kaggle”. A Bayesian logistic regression approach was utilized to incorporate prior knowledge from this historical data into the Bayesian design framework. The Bayesian optimal experimental design was then obtained by maximizing the expected utility function, with the Kullback–Leibler divergence selected as the utility criterion to support precise estimation of model parameters. The resulting optimal experiment set represents the selection of an optimal covariate set aimed at improving the prediction of the dependent variable. To determine the optimal design, the Approximate Coordinate Exchange (ACE) optimization algorithm was employed, starting with three randomly generated initial designs. This study demonstrates the efficacy of Bayesian regression analysis and optimization techniques in identifying and prioritizing key drivers impacting inventory management. By incorporating prior knowledge and accounting for uncertainty, businesses can achieve more accurate demand forecasting, improved inventory turnover, and reduced holding costs, ultimately driving superior supply chain efficiency and organizational performance.
Many statistical procedures require the assumption that the observations in a random sample are drawn from a normal distribution. Several statistical techniques, mostly based on either population moments or empirical distribution functions, are currently available to test whether the observations in a random sample are normally distributed. In this study, we use the Edge-worth series expansion to model deviations from normality, depending on the skewness and excess kurtosis of a normal population. We develop a test statistic based on these two measures to test for normality. However, the sampling distribution of the test statistic is not mathematically tractable. There-fore, we conducted a simulation study by generating the sampling distribution of the proposed test statistic for different sample sizes when the data are normally distributed. The critical values were also calculated at different levels of significance. The power of the proposed test was empirically compared with the Shapiro-Wilk (SW), Shapiro-Francia (SF), Jarque-Bera (JB), Cramer-von Mises (CM), Lilliefors (LF), Kolmogorov-Smirnov (KS), and Anderson-Darling (AD) tests. The proposed test demonstrates competitive power against JB, CM, LF, KS and AD tests for small samples under specific alternatives with low skewness.
This paper presents a comprehensive survey of recent advancements in cross-modality imaging techniques. It focuses on the application of Convolutional Neural Networks (CNNs) and Generative Adversarial Networks (GANs) in medical imaging. Cross-modality imaging involves translating images from one modality to another (E.g.: MRI to CT scans) and can assist medical experts in accurate diagnostics and patient care. CNNs and GANs have been quite popular in various fields and their abilities to work with images and generate high quality outputs are quite an advantage for a situation like this. This survey highlights several studies utilizing CNN methodologies for various applications, including synthetic MRI generation, MRI to CT translations for different anatomical regions and bone structure identification. GANs, on the other hand, excel in generative modeling by training two neural networks, the generator and the discriminator, in a competitive setting. We discuss their application in generating missing MRI modalities, creating 3D medical images, and producing synthetic CT scans for radiotherapy. The integration of CNN and GAN models in cross-modality imaging shows significant promise in improving image quality, reducing noise, and enhancing the accuracy of diagnostic tools. This survey underscores the importance of advanced deep learning techniques in medical imaging and sets the stage for future research in this rapidly evolving field.
Mixtures of chemical contaminants can pose significant health risks to humans and wildlife, even at levels deemed safe for individuals. Zebrafish (Danio rerio) offer a high-throughput exposure model with the complexity of a vertebrate organism, making them well-suited for evaluating mixture toxicity. However, substantial gaps exist in statistical methods for assessing the associations between chemical mixtures and phenotypic end-points in toxicity testing. Here, a Weighted Gene Co-expression Network Analysis (WGCNA) was developed to address this challenge, focusing on a larval behavioural assay and leveraging data from 92 well-water samples from Maine and New Hampshire, USA. Our study aims to implement mixture-relevant statistical approaches to elucidate the relationships be-tween chemicals in drinking water and the behavioural responses observed in zebrafish. A WGCNA was employed to uncover gene expression patterns that mediate behavioural effects induced by these chemical mixtures. Individual chemicals within these mixture models exhibit both positive and negative partial effects. For instance, within the turquoise module, Cadmium (Cd) is positively weighted, while Copper (Cu), Nickel (Ni), and Lead (Pb) are negatively associated with zebrafish behaviour. Similarly, within the grey module, Arsenic (As), Uranium (U) and Chromium (Cr) show positive effects, whereas Selenium (Se) and Antimony (Sb) exhibit negative effects. These findings underscore the importance of evaluating overall mixture exposure effects and highlight the critical need to consider complex interactions within chemical mixtures in environmental toxicology.
Academic libraries are essential in the modern technology environment as they provide a wide range of information and facilities through their websites, contributing to the achievement of educational goals. The quality of these dig-ital platforms is growing in importance as they influence users’ experiences and academic performance. This study aims to identify the factors that are connected with library website quality and develop a model that evaluates the effect of the quality of the library website (QoW) on user satisfaction for academic achievement (AA) by comparing service quality within the library (QoLP) and non-library website quality (QoNLW) at the University of Ruhuna using Structural Equation Modeling (SEM). For this purpose, a survey was carried out among 591 undergraduates at University of Ruhuna through a structured questionnaire. This study revealed that information quality (QoI), usability quality (QoU), and service interaction quality (QoSI) are the three main factors influencing library website quality throughout factor analysis. The SEM model showed that the library website quality explained 79.9% of the variation in user satisfaction. The study found that QoSI, QoU, QoI significantly impacted user satisfaction. Subsequently, QoW and QoLP significantly influenced academic achievement, while QoNLW had a strong impact. These results can serve as a valuable reference for university library administrators to improve the services provided on the library website and ensure students’ academic achievements by helping them develop and utilize necessary skills.