In the context of longitudinal data regression modeling, individuals often have two or more response indicators, and these response indicators are typically correlated to some extent. Additionally, in the field of clinical medicine, the response indicators of longitudinal data are often ordinal. For the joint modeling of multivariate ordinal longitudinal data, methods based on mean regression (MR) are commonly used to study latent variables. However, for data with non-normal errors, MR methods often perform poorly. As an alternative to MR methods, composite quantile regression (CQR) can overcome the limitations of MR methods and provide more robust estimates. This article proposes a joint relative composite quantile regression method (joint relative CQR) for multivariate ordinal longitudinal data and investigates its application to a set of longitudinal medical datasets on dementia. Firstly, the joint relative CQR method for multivariate ordinal longitudinal data is constructed based on the pseudo composite asymmetric Laplace distribution (PCALD) and latent variable models. Secondly, the parameter estimation problem of the model is studied using MCMC algorithms. Finally, Monte Carlo simulations and a set of longitudinal medical datasets on dementia validate the effectiveness of the proposed model and method.
In clinical medical health research, individual measurements sometimes appear as a mixture of ordinal and continuous responses. There are some statistical correlations between response indicators. Regarding the joint modeling of mixed responses, the effect of a set of explanatory variables on the conditional mean of mixed responses is usually studied based on a mean regression model. However, mean regression results tend to underperform for data with non-normal errors and outliers. Quantile regression (QR) offers not only robust estimates but also the ability to analyze the impact of explanatory variables on various quantiles of the response variable. In this paper, we propose a joint QR modeling approach for mixed ordinal and continuous responses and apply it to the analysis of a set of obesity risk data. Firstly, we construct the joint QR model for mixed ordinal and continuous responses based on multivariate asymmetric Laplace distribution and a latent variable model. Secondly, we perform parameter estimation of the model using a Markov chain Monte Carlo algorithm. Finally, Monte Carlo simulation and a set of obesity risk data analysis are used to verify the validity of the proposed model and method.
Count data is a type of data derived from the number of times an event occurs per unit of time, and zero-truncated count data refers to count data without zero. Modeling for this type of data primarily utilizes a single zero-truncated distribution, such as the zero-truncated Poisson (ZTP), zero-truncated Bell (ZTBell), and zero-truncated negative binomial (ZTNB) distribution. However, except for ZTNB model, fewer relevant models involving heterogeneous count data have been studied. Therefore, in this article, we propose a zero-truncated Bell-Poisson mixture (ZTBPM) regression model based on Bell and Poisson distributions and study the parameter estimation method of this model. Monte Carlo simulation verifies that the ZTBPM model fits better than the ZTP, ZTBell, and ZTNB regression models, providing more options for modeling zero-truncated count data and solving the heterogeneity problem more effectively. Finally, the ZTBPM regression model is applied to a set of outpatient data analysis to study the risk factors affecting the number of residents' outpatient visits. This is of great significance for better understanding the health status and medical needs of residents, optimizing the allocation of medical resources, and guiding the formulation of medical policies.
Link prediction has traditionally been regarded as a binary classification problem, aiming to predict whether a link exists between two nodes in a given network. However, this binary framework fails to account for the cooperation intensity or the diversity of relationships. For example, in collaboration networks, the cooperation intensity often varies depending on the number of collaborations. Therefore, building on the premise of existing collaborations, this study models the relationships between authors as an ordinal multiclass problem to more accurately characterize varying levels of cooperation intensity. Then, the ordinal collaboration network model with zero-truncated Poisson latent variables ZTP-OCN$$ \left(\mathrm{ZTP}\hbox{-} \mathrm{OCN}\right) $$ is constructed. The maximum likelihood estimation MLE$$ \left(\mathrm{MLE}\right) $$ method is used to estimate the model parameters, and the performance of the model is evaluated by numerical simulation. Finally, this paper applies the ZTP-OCN model to the collaboration network of statistical journals to verify its validity in predicting the cooperation intensity. The results show that the model can describe the cooperation relationship with different intensity well.
With the quick development of society and industry, air quality has become a grim and global environmental concern. Predicting and rating air quality for many cities remains a significant challenge. Consequently, machine learning algorithms have garnered considerable attention for their potential to address these issues effectively. In this paper, firstly, based on daily air quality data from July 1, 2022 to June 30, 2023 in Lanzhou city of China, five machine learning models, including Bayes Model Averaging (BMA), Support Vector Machine (SVM), Gradient Boosting Decision Tree (GBDT), Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) are developed to predict the Air Quality Index (AQI) via six major air pollutants (PM2.5, PM10, SO2, NO2, O3 and CO). Secondly, we integrate Bootstrap algorithm into the optimal model, leading to the proposal of the LSTM-Bootstrap algorithm for deriving the standard errors and confidence intervals of the predicted AQI. Thirdly, a cumulative logit model is employed to evaluate and forecast AQI rating. The analysis results indicate that AQI rating is significantly affected by PM10, CO and O3. Additionally, to validate the efficacy of the suggested methods, a similar analysis is conducted on air quality data from Chengdu city for the same period. The findings provide valuable insights for future environmental policies and air quality management strategies.
The rapid expansion of global tourism has underscored the critical need for accurate regional tourism revenue (TR) forecasting to support sustainable economic development. This study takes Shaanxi Province, a major tourism destination in China, as a case to analyse TR trends and predict future changes. We develop a Quantile Neural Network (QNN) model, optimized through grid search and ten-fold cross-validation, which demonstrates superior predictive accuracy compared to benchmark models. To quantify prediction uncertainty, we propose the QNN-Bootstrap algorithm that combines QNN with resampling to construct confidence intervals, enhancing forecast reliability. Furthermore, the model's generalizability is validated using tourism data from Beijing, confirming its robust performance in diverse regional contexts. To assess potential risks, this study also simulates TR dynamics under external shocks such as economic crises, pandemics, and natural disasters, and discusses the recovery trajectory following the COVID-19 pandemic. The findings provide valuable insights and policy recommendations to support evidence-based decision-making and promote resilient and sustainable tourism development.
This paper presents the joint parameters inference of conditional quantiles for a multivariate response linear regression model with a vector autoregressive (VAR) error using the expectation-maximization (EM) algorithm. Because the error follows a VAR model, the proposed approach accounts for the associations among multivariate responses and how the relationships between responses and explanatory variables vary across different quantiles of the marginal conditional distribution of responses. To facilitate likelihood-based inference using the EM algorithm, a multivariate asymmetric Laplace (MAL) distribution is forced on the independent errors of the model, thereby allowing the construction of an equivalently joint quantile model. Meanwhile, a location-scale mixture representation of the MAL distribution is employed to simplify the model’s working likelihood structure. Last, we present simulation studies and the analysis of real data for concerning on energy efficiency evaluation in order to illustrate the proposed modeling approach’s effectiveness.
In 2013, China launched a cooperative initiative to build the "Silk Road Economic Belt" and the "21st Century Maritime Silk Road". Based on the data of import and export trade between 57 countries along the "Belt and Road" and China from 2010 to 2021, this paper conducts research using the spatial Durbin model. The results show that the double fixed-effects model is optimal. The 57 countries along the "Belt and Road" have significant positive spatial spillover effects on trade with China, among which economic development, educational development, and transportation accessibility have significantly promoted trade between the countries along the "Belt and Road" and China.
With the acceleration of economic development and urbanization, air pollution has become increasingly severe and has been a crucial issue affecting social advancement. Considering the spatial correlation between regions in air quality analysis can improve the accuracy of model estimation for the data on air pollution. First, we propose the functional Logistic regression model with spatial effects. Second, we fit the original data into functional data using B-spline basis functions and apply functional principal component analysis for dimension reduction. Further, the model is estimated using the maximum likelihood method. Finally, the effectiveness of the proposed model is validated through numerical simulations and a real data analysis for PM2.5 air quality in the Sichuan-Chongqing region of China.
With the growing application of collaboration networks in scientific research, link prediction has become an essential tool for understanding the structural and evolutionary mechanisms of academic relationships. Traditional link prediction methods typically model links as a binary classification problem, predicting whether a link exists between a pair of nodes. However, in real-world networks, collaboration intensity varies, and simple 0-1 classification fails to capture such heterogeneity. To address this limitation, we propose a longitudinal ordinal collaboration network model for multiclass link prediction, which accounts for both the ordinal nature of collaboration intensity and the temporal evolution of collaborative relationships. Under the maximum likelihood framework, an regularization term is incorporated to enhance the stability and interpretability of parameter estimation. The proposed model extends link prediction from binary to ordinal multiclass classification and explicitly introduces a longitudinal structure to characterize the evolution of collaborations over time. Simulation studies and empirical analyses demonstrate that the model offers strong explanatory power and predictive performance, providing an effective tool for dynamic modelling of collaboration networks.
ABSTRACTIn the era of big data, detecting outliers in time series data is crucial, particularly in fields such as finance and engineering. This article proposes a novel sequence outlier detection method based on the gated recurrent unit autoencoder with Gaussian mixture model (GRU‐AE‐GMM), which combines gated recurrent unit (GRU), autoencoder (AE), Gaussian mixture model (GMM), and optimization algorithms. The GRU captures long‐term dependencies within the sequence, while the AE measures sequence abnormality. Meanwhile, the GMM models the relationship between the original and reconstructed sequences, employing the Expectation–Maximization (EM) algorithm for parameter estimation to calculate the likelihood of each hidden variable belonging to each Gaussian mixture component. In this article, we first train the model with mean‐squared error loss (MSEL), and then further enhanced by substituting it with quantile loss (QL), composite quantile loss (CQL), and Huber loss (HL), respectively. Next, we validate the effectiveness and robustness of the proposed model through Monte Carlo experiments conducted under different error terms. Finally, the method is applied to Amazon stock data for 2022, demonstrating its significant potential for application in dynamic and unpredictable market environments.
In this paper, we propose a Bayesian joint quantile regression approach to model multivariate ordinal longitudinal data with multi-response mixed effects model. A multivariate asymmetric Laplace (MAL) distribution is employed to construct the working likelihood of the considered model. LASSO-type penalization priors of regression parameters are incorporated into the working likelihood to implement Bayesian joint QR inference. MCMC algorithm is employed to derive the fully posterior distributions of all parameters. For multivariate ordinal longitudinal data, in order to give more efficient estimation results, we propose the joint relatively QR approach for regression coefficients. Some simulation studies are presented to illustrate the performance of the proposed joint relatively QR estimation approach. Finally, we analyze a multivariate ordinal longitudinal dataset from a clinical study in diagnosis and treatment of lung cancer using the proposed approach.
In the era of big data, accurately predicting trends and uncertainties in time series data is crucial for various fields, such as finance and engineering. This paper proposes an adaptive weighted regularized quantile regression gated recurrent unit (AWR-QRGRU) algorithm. The gated recurrent unit (GRU) is employed to capture long-term dependencies in the sequence and generate predictions for multiple quantiles through a fully connected layer, thereby enabling both point forecasts and interval forecasts with varying confidence levels. In the proposed prediction algorithm, the loss function incorporates coverage loss, median loss, interval width loss, and regularization loss. An adaptive weight adjustment mechanism was implemented to dynamically optimize the weights of these different loss terms, enhancing the accuracy and stability of the predictions when dealing with multidimensional explanatory variables. Subsequently, we conducted Monte Carlo experiments to validate the algorithm's effectiveness in both point and interval predictions. Ultimately, the algorithm was applied to predict Amazon's and NVIDIA's stock prices, showcasing its potential applications in complex financial market environments.
In longitudinal data regression modeling, an individual measurements often include two or more responses. Especially in the field of clinical health research, there is usually some correlation between the response variables in longitudinal data. For joint modeling of multi-response longitudinal data, research is typically based on mean regression (MR) methods. However, for data with non-normal errors and the presence of outliers, mean regression (MR) methods often perform poorly. Composite quantile regression (CQR) methods not only provide robust estimates but also allow for the examination of the combined effects of a set of explanatory variables on multiple quantile points of the response variable. This paper proposes a joint composite quantile regression method for multi-response longitudinal data and investigates its application in a set of longitudinal medical datasets related to liver cirrhosis. First, a joint CQR method for multi-response longitudinal data is constructed based on the Pseudo Composite Asymmetric Laplace Distribution (PCALD) and latent variable models. Next, the MCMC algorithm is used to investigate the parameter estimation issues of the model. Finally, the effectiveness of the proposed model and methods is validated through Monte Carlo simulations and analysis of a set of longitudinal medical datasets related to liver cirrhosis.
In real applied fields such as clinical medicine, environmental sciences, psychology as well as economics, we often encounter the task of conducting statistical inference for longitudinal data with ordinal responses. The traditional methods of longitudinal data analysis are often inclined to model continuous responses, which are no longer suitable for such ordinal data. Logistic regression and probit regression are two considerable methods which are frequently used to model ordinal longitudinal responses. However, such modelling methods just depict the mean feature of latent outcome variable and may produce non-robust results when encountering nor-normal errors or outliers. As a proper alternative of mean regression models, composite quantile regression (CQR) method is usually employed to derive robust estimation. The target of this paper is to investigate the CQR estimation approach for ordinal latent longitudinal model. The joint Bayesian hierarchical model is established and a relative CQR estimation approach is suggested to conduct posterior inference for the considered model. Further, in longitudinal data modelling, excessive predictors may be brought into in the models which result in the decrease of the model prediction precision. Bayesian L1/2 regularized prior is incorporated into ordinal longitudinal CQR model to conduct variable selection simultaneously. Finally, simulation studies and two ordinal longitudinal data analysis are hired to illustrate the considered method.
In this paper, we propose a Bayesian quantile regression (QR) approach to jointly model multivariate ordinal data. Firstly, a multivariate latent variable model is used to link the multivariate ordinal data and latent continuous responses and the multivariate asymmetric Laplace (MAL) distribution is employed to construct the joint QR-based working likelihood for the considered model. Secondly, adaptive- L_1/2 penalization priors of regression parameters are incorporated into the working likelihood to implement high-dimensional Bayesian joint QR inference. Markov Chain Monte Carlo (MCMC) algorithm is utilized to derive the fully conditional posterior distributions of all parameters. Thirdly, Bayesian joint relatively QR estimation approach is recommended to result in more efficient estimation results. Finally, Monte Carlo simulation studies and a real instance analysis of multirater agreement data are presented to illustrate the performance of the proposed Bayesian joint relatively QR approach.
Along with the rapid development of industries and the acceleration of urbanisation, the problem of air pollution is becoming more serious. Exploring the relevant factors affecting air quality and accurately predicting the air quality index are significant in improving the overall environmental quality and realising green economic development. Machine learning algorithms and statistical models have been widely used in air quality prediction and ranking assessment. In this paper, based on daily air quality data for the city of Xi’an, China, from 1 October 2022 to 30 September 2023, we construct support vector regression (SVR), gradient boosting decision tree (GBDT), extreme gradient boosting (XGBoost), random forests (RF), neural network (NN) and long short-term memory (LSTM) models to analyse the influence of the air quality index for Xi’an and to conduct comparative tests. The predicted values and 95% prediction intervals of the AQI for the next 15 days for Xi’an, China, are given based on the Bootstrap-XGBoost algorithm. Further, the ordinal logit regression and ordinal probit regression models are constructed to evaluate and accurately predict the AQI ranks of the data from 1 October 2023 to 15 October 2023 for Xi’an. Finally, this paper proposes some suggestions and policy measures based on the findings of this paper.
Agricultural and rural carbon (ARC) emissions are a major source of greenhouse gas emissions in China and have profound implications for implementing the rural revitalization strategy. This study takes Shandong Province, a leading agricultural province in China, as a case study to explore the relationship between ARC emissions and their influencing factors. It employs the Logarithmic Mean Divisia Index (LMDI) model to decompose changes in ARC emissions from 2000 to 2021, analyzing the contributions of factors such as agricultural production efficiency and agricultural industrial structure. The study then expands the indicator system and applies feature selection methods to identify the main influencing factors. It establishes Bayes model averaging (BMA), STIRPAT-Ridge regression and Long Short-Term Memory (LSTM) models to evaluate their performance in modeling historical ARC emissions. Finally, the study makes prospective forecasts of ARC emissions in Shandong Province from 2022 to 2050 under low, medium and high speed development scenarios. The findings show that from 2000 to 2021, ARC emission intensity decreased by 71.86
In practical data analysis, individual measurements usually include two or more responses, and some statistical correlations often exist between the responses. Especially in medical data analysis, observations are often binary responses. A class of multi-response logistic regression model based on a joint modeling approach is investigated in this paper, and an application to a group data of primary biliary cirrhosis diseases is considered. Firstly, we propose a new class of multi-response logistic distribution and investigate its statistical properties. Secondly, a multi-response logistic regression model is constructed using a latent variable model and multi-variate logistic error distribution. Furthermore, the parameter estimation method of the model is provided by applying the monte carlo expectation maximization (MCEM) algorithm and the multiple imputation method. Finally, numerical simulations and comparative predictions on a test set are performed to validate the finite sample performance of the proposed model, and the model is applied to a cirrhosis disease dataset for analysis.
Ordinal data frequently occur in various fields such as knowledge level assessment, credit rating, clinical disease diagnosis, and psychological evaluation. The classic models including cumulative logistic regression or probit regression are often used to model such ordinal data. But these modeling approaches conditionally depict the mean characteristic of response variable on a cluster of predictive variables, which often results in non-robust estimation results. As a considerable alternative, composite quantile regression (CQR) approach is usually employed to gain more robust and relatively efficient results. In this paper, we propose a Bayesian CQR modeling approach for ordinal latent regression model. In order to overcome the recognizability problem of the considered model and obtain more robust estimation results, we advocate to using the Bayesian relative CQR approach to estimate regression parameters. Additionally, in regression modeling, it is a highly desirable task to obtain a parsimonious model that retains only important covariates. We incorporate the Bayesian L-1/2 penalty into the ordinal latent CQR regression model to simultaneously conduct parameter estimation and variable selection. Finally, the proposed Bayesian relative CQR approach is illustrated by Monte Carlo simulations and a real data application. Simulation results and real data examples show that the suggested Bayesian relative CQR approach has good performance for the ordinal regression models.