Background & Aims: Hepatocellular carcinoma (HCC) is characterized by a high mortality rate. The Liver Imaging Reporting and Data System (LI-RADS) results in a considerable number of indeterminate observations, rendering an accurate diagnosis difficult. Methods: We developed four deep learning models for diagnosing HCC on computed tomography (CT) via a training-validation- testing approach. Thin-slice triphasic CT liver images and relevant clinical information were collected and processed for deep learning. HCC was diagnosed and verified via a 12-month clinical composite reference standard. CT observations among at-risk patients were annotated using LI-RADS. Diagnostic performance was assessed by internal validation and independent external testing. We conducted sensitivity analyses of different subgroups, deep learning explainability evaluation, and misclassification analysis. Results: From 2,832 patients and 4,305 CT observations, the best-performing model was Spatio-Temporal 3D Convolution Network (ST3DCN), achieving area under receiver-operating-characteristic curves (AUCs) of 0.919 (95% CI, 0.903-0.935) and 0.901 (95% CI, 0.879-0.924) at the observation (n = 1,077) and patient (n = 685) levels, respectively during internal validation, compared with 0.839 (95% CI, 0.814-0.864) and 0.822 (95% CI, 0.790-0.853), respectively for standard of care radiological interpretation. The negative predictive values of ST3DCN were 0.966 (95% CI, 0.954-0.979) and 0.951 (95% CI, 0.931-0.971), respectively. The observation-level AUCs among at-risk patients, 2-5-cm observations, and singular portovenous phase analysis of ST3DCN were 0.899 (95% CI, 0.874-0.924), 0.872 (95% CI, 0.838-0.909) and 0.912 (95% CI, 0.895-0.929), respectively. In external testing (551/717 patients/observations), the AUC of ST3DCN was 0.901 (95% CI, 0.877-0.924), which was non-inferior to radiological interpretation (AUC 0.900; 95% CI, 0.877--923). Conclusions: ST3DCN achieved strong, robust performance for accurate HCC diagnosis on CT. Thus, deep learning can expedite and improve the process of diagnosing HCC. (c) 2024 The Author(s). Published by Elsevier B.V. on behalf of European Association for the Study of the Liver (EASL). This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
There has been growing interest in extending the popular threshold time series models to include a buffer zone for regime transition. However, almost all attention has been on buffered autoregressive models. Note that the classical moving average (MA) model plays an equally important role as the autoregressive model in classical time series analysis. It is therefore natural to extend our investigation to the buffered MA (BMA) model. We focus on the first-order BMA model while extending to more general MA model should be direct in principle. The proposed model shares the piecewise linear structure of the threshold model, but has a more flexible regime switching mechanism. Its probabilistic structure is studied to some extent. A nonlinear least squares estimation procedure is proposed. Under some standard regularity conditions, the estimator is strongly consistent and the estimator of the coefficients is asymptotically normal when the parameter of the boundary of the buffer zone is known. A portmanteau goodness-of-fit test is derived. Simulation results and empirical examples are carried out and lend further support to the usefulness of the BMA model and the asymptotic results.
This paper derives the asymptotic distribution of the least absolute deviations estimator for nonstationary vector autoregressive time series models with pure unit roots under mild conditions. As this distribution has a complicated form, many commonly used bootstrap techniques can-not be directly applied. To tackle this problem, we propose a novel hybrid bootstrap method by combining the classical wild bootstrap and the method in [17]. We establish the asymptotic validity of the proposed method and further apply it to construct three bootstrapping panel unit root tests. Monte Carlo experiments support the validity of our inference procedure in finite samples. The usefulness of the proposed panel unit root tests is demonstrated via analyses of real economic and financial data sets.
We first construct a new generalized Hausman test for detecting the structural change in a multiplicative form of covariance matrix time series model. This generalized Hausman test is asymptotically pivotal, and has nontrivial power in detecting a broad class of alternatives. Moreover, we propose a new semiparametric covariance matrix time series model. The proposed model has a time-varying longrun component that takes the structural change into account, and a BEKK-type short-run component that captures the temporal dependence. We propose a twostep estimation procedure to estimate this semiparametric model, and establish the asymptotics of the related estimators. Finally, the importance of the generalized Hausman test and the semiparametric model is illustrated by means of simulations and an application to realized covariance matrix data.
Asymmetric power GARCH models have been widely used to study the higher order moments of financial returns, while their quantile estimation has been rarely investigated. This paper introduces a simple monotonic transformation on its conditional quantile function to make the quantile regression tractable. The asymptotic normality of the resulting quantile estimators is established under either stationarity or non-stationarity. Moreover, based on the estimation procedure, new tests for strict stationarity and asymmetry are also constructed. This is the first try of the quantile estimation for non-stationary ARCH-type models in the literature. The usefulness of the proposed methodology is illustrated by simulation results and real data analysis.
We propose a new Conditional BEKK matrix-F (CBF) model for the time-varying realized covariance (RCOV) matrices. This CBF model is capable of capturing heavy-tailed RCOV, which is an important stylized fact but could not be handled adequately by the Wishart-based models. To further mimic the long memory feature of the RCOV, a special CBF model with the conditional heterogeneous autoregressive (HAR) structure is introduced. Moreover, we give a systematical study on the probabilistic properties and statistical inferences of the CBF model, including exploring its stationarity, establishing the asymptotics of its maximum likelihood estimator, and giving some new inner-product-based tests for its model checking. In order to handle a large dimensional RCOV matrix, we construct two reduced CBF models -- the variance-target CBF model (for moderate but fixed dimensional RCOV matrix) and the factor CBF model (for high dimensional RCOV matrix). For both reduced models, the asymptotic theory of the estimated parameters is derived. The importance of our entire methodology is illustrated by simulation results and two real examples.
This paper considers the mixture autoregressive panel (MARP) model. This model can capture the burst and multi-modal phenomenon in some panel data sets. It also enlarges the stationarity region of the traditional AR model. An estimation method based on the EM algorithm is proposed and the assumption required of the model is quite low. To illustrate the method, we fitted the MARP model to the gray-sided voles data. Another MARP model with less restriction is also proposed.
The proposition of tail risk as a new asset pricing factor has gained traction in recent years. Recent work by Almeida, Ardison, Garcia, and Vicente (Nonparametric tail risk, stock returns, and the macroeconomy.J. Financ. Economet., 2017,15(3), 333-376) proxies the cross-sectional variation in returns by Fama-French portfolios, which are further summarized into a few basis assets via principal component analysis. The number of states of nature is set higher than that of the basis assets to estimate the stochastic discount factors, which in turn risk-neutralize the excess expected shortfall as a tail risk measure. As an alternative approach to this means of dimension reduction, we propose tackling the problem directly by forming portfolios that minimize the excess expected shortfall. Our direct measure exhibits greater explanatory power when applied to more liquid, nonlottery-style stock returns. More importantly, our proposed approach reveals direct exposure to downside systematic risk without the need for an additional risk-neutralization step.
This paper proposes some novel one-sided omnibus tests for independence between two multivariate stationary time series. These new tests apply the Hilbert-Schmidt independence criterion (HSIC) to test the independence between the innovations of both time series. Under regular conditions, the limiting null distributions of our HSIC-based tests are established. Next, our HSIC-based tests are shown to be consistent. Moreover, a residual bootstrap method is used to obtain the critical values for our HSIC-based tests, and its validity is justified. Compared with the existing cross-correlation-based tests for linear dependence, our tests examine the general (including both linear and non-linear) dependence to give investigators more complete information on the causal relationship between two multivariate time series. The merits of our tests are illustrated by some simulation results and a real example.
The conditional approach of Hill's estimator depends on a threshold choice. This may give different results in a statistical test when different thresholds are used. Motivated by the uniformly most powerful test, this article proposes a new test for heavy tail of various degrees. Simulation shows that the test is applicable and its power is superior to three existing methods in the literature. An example of the Danish fire loss data elucidates the inconsistent conclusion in existing statistical tests and demonstrates the consistent conclusion the proposed test leads to. A new and affirmative insight into the data is that the variance does not exist.
Modelling and forecasting covariance matrices of asset returns play a crucial role in many financial fields, such as portfolio allocation and asset pricing. The availability of high-frequency intraday data enables the modelling of the realized covariance matrix directly. However, most models in the literature suffer from the curse of dimensionality, i.e. the number of parameters needed increases at the rate of the square of the number of assets. To solve the problem, we propose a factor model with a diagonal Conditional Autoregressive Wishart model for the factor realized covariance matrices. Consequently, the positive definiteness of the estimated covariance matrix is ensured with the proposed model. Asymptotic theory is derived for the estimated parameters. In the extensive empirical analysis, we find that the number of parameters can be reduced significantly; to only about one-tenth of the benchmark model. Furthermore, the proposed model maintains a comparable performance with a benchmark vector autoregressive model for different forecast horizons.
A fuzzy portfolio selection model is considered with a view to incorporating ambiguity about model and data structure. The model features the uncertainty about the exit time of each risky asset within a pre-specified investment horizon and also the presence of transaction costs. However, departing from the traditional paradigm where the transaction costs are often assumed to be unrelated to holding periods, we introduce the capital gain tax of which the realized tax rate is decreasing with respect to the holding periods with a view to encouraging the long-term investment. Meanwhile, the regime switching property of the market state is introduced to fuzzy portfolio selection, where fuzzy random variables are employed to model uncertain returns of risky assets in a Markov-regime switching market. An adjusted L - R fuzzy number is introduced and some of its mathematical properties are studied. In addition, a bi-objective mean-variance model is formulated, and a time varying numerical integral-based particle swarm optimization algorithm (TVNIPSO) is designed to obtain the efficient frontier of the portfolio in the sense of Pareto dominance. Finally, some numerical experiments are provided to validate the effectiveness of the model and the TVNIPSO. (C) 2020 Elsevier Ltd. All rights reserved.
A buffered autoregression extends the classical threshold autoregression by allowing a buffer region for regime changes. In this study, we examine asymptotic statistical inferences for the two-regime buffered autoregressive (BAR) model, with autoregressive unit roots. We propose a Sup-LR test for the nonlinear buffer effect in the possible presence of unit roots, and a class of unit root tests to identify the number of nonstationary regimes in the BAR model. The wild bootstrap method is suggested to approximate the critical values of the two tests. Simulation results show that the proposed unit root test outperforms the conventional augmented Dickey-Fuller test, and that the two wild bootstrap tests are robust to unknown heteroscedasticity. Two macroeconomic data examples, based on U.S. unemployment rates and real exchange rates, respectively, are provided to illustrate the methods.
This article investigates a portmanteau test statistic for checking model adequacy of smooth transition autoregressive (STAR) models. The asymptotic distribution of residual autocorrelations and the least‐squares estimators are also derived. Hence, the correct asymptotic standard errors for residual autocorrelations are also obtained facilitating model diagnostic checking. Through the graphical display of the simulation results concerning the size and power, for commonly used nominal sizes (), the portmanteau test appears to be more advantageous than the Lagrange multiplier tests in checking serial independence for the errors of STAR models.
Variable screening for censored survival data is most challenging when both survival and censoring times are correlated with an ultrahigh-dimensional vector of covariates. Existing approaches to handling censoring often make use of inverse probability weighting by assuming independent censoring with both survival time and covariates. This is a convenient but rather restrictive assumption which may be unmet in real applications, especially when the censoring mechanism is complex and the number of covariates is large. To accommodate heterogeneous (covariate-dependent) censoring that is often present in high-dimensional survival data, we propose a Gehan-type rank screening method to select features that are relevant to the survival time. The method is invariant to monotone transformations of the response and of the predictors, and works robustly for a general class of survival models. We establish the sure screening property of the proposed methodology. Simulation studies and a lymphoma data analysis demonstrate its favorable performance and practical utility.
This article explores the fitting of Autoregressive (AR) and Threshold AR (TAR) models with a non-Gaussian error structure. This is motivated by the problem of finding a possible probabilistic model for the realized volatility. A Gamma random error is proposed to cater for the non-negativity of the realized volatility. With many good properties, such as consistency even for non-Gaussian errors, the maximum likelihood estimate is applied. Furthermore, a non-gradient numerical Nelder–Mead method for optimization and a penalty method, introduced for the non-negative constraint imposed by the Gamma distribution, are used. In the simulation experiments, the proposed fitting method found the true model with a rather insignificant bias and mean square error (MSE), given the true AR or TAR model. The AR and TAR models with Gamma random error are then tested on empirical realized volatility data of 30 stocks, where one third of the cases are fitted quite well, suggesting that the model may have potential as a supplement for current Gaussian random error models with proper adaptation.
To see the impact of the kernel functions k and l, we examine the performance of our HSIC-based test statistics when k and l are chosen as inverse multi-quadratics kernels with α = β = 1. Tables S1-S2 report the sizes and power of all examined HSIC-based tests. Compared with the results in Tables 1-2, the results in Tables S1-S2 imply that similar performance of our HSIC-based tests retains for the two different choices of kernels in most cases, but some differences may exist in some cases. This is consistent with the findings in Gretton et al. (2009).
Recently, inference about high-dimensional integrated covariance matrices (ICVs) based on noisy high-frequency data has emerged as a challenging problem. In the literature, a pre-averaging estimator (PA-RCov) is proposed to deal with the microstructure noise. Using the large-dimensional random matrix theory, it has been established that the eigenvalue distribution of the PA-RCov matrix is intimately linked to that of the ICV through the Marčenko–Pasturequation. Consequently, the spectrum of the ICV can be inferred from that of the PA-RCov. However, extensive data analyses demonstrate that the spectrum of the PA-RCov is spiked, that is, a few large eigenvalues (spikes) stay away from the others which form a rather continuous distribution with a density function (bulk). Therefore, any inference on the ICVs must take into account this spiked structure. As a methodological contribution, a spiked model is proposed for the ICVs where spikes can be inferred from those of the available PA-RCov matrices. The consistency of the inference procedure is established. In addition, the methodology is applied to the real data from the US and Hong Kong markets. It is found that the model clearly outperforms the existing one in predicting the existence of spikes and in mimicking the empirical PA-RCov matrices.