
Commonly used variance estimators for the generalized regression estimator (GREG) are based on Taylor linearization and jackknife. Traditionally, a jackknife GREG variance estimator is obtained by jackknifing GREG, which consists of computing GREG from each of several subsamples of the parent sample, and estimating the variance of the parent sample from the variability between the subsample GREG estimates. The jackknife technique can also be used for bias correction, which means that the jackknifing of GREG can be used to reduce the bias of GREG. Instead of jackknifing GREG, we propose to jackknife a variance estimator of GREG to obtain a bias-corrected GREG variance estimator. We show that jackknifing a variance estimator can be performed as usual by viewing the variance estimator as the mean over a sample selected in a synthetic finite population with a sampling design having first-order inclusion probabilities equal to the second-order inclusion probabilities of the initial sampling design. Les estimateurs de variance couramment utilis & eacute;s pour l'estimateur par r & eacute;gression g & eacute;n & eacute;ralis & eacute;e (GREG) reposent sur la lin & eacute;arisation de Taylor et la m & eacute;thode du jackknife. Dans sa mise en oe uvre classique, le jackknife consiste & agrave; recalculer le GREG sur des sous-& eacute;chantillons issus de l'& eacute;chantillon initial, puis & agrave; estimer la variance de l'& eacute;chantillon parent & agrave; partir de la variabilit & eacute; entre ces estimations partielles. La m & eacute;thode du jackknife peut & eacute;galement & ecirc;tre employ & eacute;e & agrave; des fins de correction du biais. Appliqu & eacute;e au GREG, elle vise alors & agrave; r & eacute;duire le biais propre & agrave; cet estimateur. La d & eacute;marche propos & eacute;e ici s'inscrit dans une perspective diff & eacute;rente. Plut & ocirc;t que d'appliquer le jackknife directement & agrave; l'estimateur GREG, nous l'appliquons & agrave; un estimateur de sa variance, dans l'objectif d'obtenir un estimateur de variance du GREG avec correction du biais. Nous montrons que cette proc & eacute;dure peut suivre le sch & eacute;ma usuel, en interpr & eacute;tant l'estimateur de variance comme la moyenne d'un & eacute;chantillon tir & eacute; d'une population finie synth & eacute;tique. Le plan de sondage associ & eacute; est d & eacute;fini de sorte que les probabilit & eacute;s d'inclusion du premier ordre co & iuml;ncident avec les probabilit & eacute;s d'inclusion du second ordre du plan initial.
Uncertainty quantification in functional data classification is crucial for reliable decision-making but remains insufficiently addressed in existing methods, which predominantly focus on point predictions. In this paper we propose a novel method for constructing prediction sets within the conformal prediction framework. Our approach begins by improving the construction of functional prediction bands, enhancing their ability to capture the shape characteristics of functional data. Based on these refined bands, we introduce a new nonconformity measure that quantifies the signed distance between a curve and the class-specific prediction bands. This enables us to construct label prediction sets with valid coverage guarantees at a given confidence level. Extensive simulation studies and real data experiments demonstrate that the proposed method yields prediction sets with high coverage and small ambiguity, effectively quantifying uncertainty while maintaining strong classification performance.
A novel algorithm combining both node degree and edge betweenness (node degree+edge betweenness) is developed to highlight gradients in connectivity. It identifies communities within complex directed graphs. These are networks that are composed of actors who have ties that originate and end at themselves (self-loops). Hierarchies in social partitioning for node degree (in-ties, out-ties) are evident among actors when applying the node degree+edge betweenness algorithm to network datasets. This observed gradient characteristic has useful application in epidemiological and health services research where the direction and temporality of ties matter in activities such as contact tracing and the deployment of resources. Our new algorithm is readily available in the R package ig.degree.betweenness and in the Python module ig-degree-betweenness.
In the new era of big data, modelling multivariate spatial-temporal data is a challenging task due to both the high dimensionality of the features and complex associations among the responses across different locations and time points. To improve the estimation efficiency, we propose a spatial-temporal partial envelope model that is parsimonious and effective in modelling high-dimensional spatial-temporal data. The partial envelope model is proposed under a linear coregionalization model framework, which allows heterogeneous covariance structures for different variables of the response vector. We study the asymptotic behaviour of the estimator and conduct a thorough simulation study to demonstrate the soundness and effectiveness of the proposed method. We also apply the proposed model to analyze the crowdsourcing weather data collected from personal weather stations in the city of Syracuse, New York, USA. Dans la nouvelle & egrave;re des m & eacute;gadonn & eacute;es, la mod & eacute;lisation de donn & eacute;es spatio-temporelles multivari & eacute;es constitue une t & acirc;che difficile en raison & agrave; la fois de la haute dimensionnalit & eacute; des variables explicatives et des associations complexes entre les r & eacute;ponses, observ & eacute;es & agrave; diff & eacute;rents emplacements et moments. Pour am & eacute;liorer l'efficacit & eacute; de l'estimation, nous proposons un mod & egrave;le d'enveloppe partielle spatio-temporelle, parcimonieux et efficace pour mod & eacute;liser des donn & eacute;es spatio-temporelles de haute dimension. Ce mod & egrave;le est d & eacute;velopp & eacute; dans le cadre d'un mod & egrave;le de co-r & eacute;gionalisation lin & eacute;aire, qui autorise des structures de covariance h & eacute;t & eacute;rog & egrave;nes pour les diff & eacute;rentes variables du vecteur de r & eacute;ponse. Nous & eacute;tudions le comportement asymptotique de l'estimateur et r & eacute;alisons une & eacute;tude de simulation approfondie afin de d & eacute;montrer la validit & eacute; et l'efficacit & eacute; de la m & eacute;thode propos & eacute;e. Nous appliquons & eacute;galement ce mod & egrave;le & agrave; l'analyse de donn & eacute;es m & eacute;t & eacute;orologiques issues de l'externalisation ouverte, collect & eacute;es par des stations m & eacute;t & eacute;orologiques personnelles dans la ville de Syracuse, dans l'& Eacute;tat de New York, aux & Eacute;tats-Unis.
When used to estimate variance components (VCs), confidence intervals (CIs) can be truncated at zero, have a point estimate not in the quoted CI, be empty with positive probability, or be all-inclusive. This is because they have conflicting dual roles, since they are considered to cover the parameter with a specified probability while also indicating the precision of an estimate. Such results provide counterexamples to the belief that CIs indicate precise estimation and are a validating principle. Repetitions with similar results do not provide a scientific proof-which is disproven by a single counterexample. Our Gaussian theory of error, initiated by J.W. Tukey, only indicates the precision of an estimate, and provides a useful estimate for a unique sample. Simulated distributions of propagated error and margins of error can estimate VCs and non-negative functions of VCs in balanced random effects models, and in one-way unbalanced random effects models, providing useful estimates. Decisions made after Gaussian estimation are independent acts.
Variations in strains of COVID-19 have a significant impact on the rate of surges and on the accuracy of forecasts of the epidemic dynamics. The primary goal for this article is to quantify the effects of varying strains of COVID-19 on ensemble forecasts of individual "surges." By modelling the disease dynamics with an SIR model, we solve the inverse ensemble forecasting problem to compute a probability distribution on the transmission parameters from observed population data at an early specified time during a surge. Using the computed distribution, we then forecast future surge dynamics. The solution of the inverse ensemble forecasting problem is computed by applying a Bayesian approach to the disintegration of measures. We verify the method and explore properties of the inverse solution using synthetic data. We also make ensemble forecasts using real data obtained from a specific population of towns.
Intensity-duration-frequency curves are used by a wide range of professionals to manage the risks related to extreme rainfall. In Canada, these curves are produced by Environment and Climate Change Canada on the basis of Gumbel distributions fitted independently for each accumulation period. Generalized extreme-value (GEV) distributions are more flexible, but uncertainties in parameter estimates can lead to physical inconsistencies across durations. The scaling GEV model offers a solution to this issue, provided that the presence of dependent maxima across different periods is accounted for. Two options are considered: (i) modelling the inter-duration dependence using an elliptical copula with a Mat & eacute;rn correlation matrix, and (ii) working from a composite likelihood assuming independence and estimating uncertainty using Godambe's information matrix. The two strategies are compared using Canadian data to guide practitioners in choosing the approach best suited to their needs.
Index insurance design involves integrating weather data, soil moisture, phenology information, and satellite imagery, which presents challenges in data fusion. This article addresses the modelling of multisource functional indices of varying lengths by constructing a stagewise ensemble of sequential models. The implemented methods, including nonparametric regression and deep learning models, aim to improve crop yield prediction by systematically capturing spatiotemporal dependence across indices of different temporal spans. Results from an applied case study demonstrate both the feasibility and practical value of stagewise modelling, highlighting its potential to reduce basis risk and improve the hedging effectiveness of index insurance contracts.
Regression models are often used to analyze discrete outcomes, but classical goodness-of-fit tests such as those based on the deviance or Pearson's statistic can be misleading or have little power in this context. To address this issue, we propose a new test, inspired by the work of Czado et al. (Biometrics, 65(4):1254-1261, 2009), which involves no randomization, tuning parameter, or binning of covariates. The statistic's large-sample distribution under the null hypothesis is determined; as it involves unknown parameter values, one must resort to a bootstrap procedure to compute -values. Simulations are conducted to investigate the ability of the test to detect a broad range of model misspecifications commonly seen in practice. The proposed procedure is seen to perform well in all the scenarios considered as well as on real data.
The asymptotic dependence structure between multivariate extreme values is fully characterized by their projections on the unit simplex. Under mild conditions, the only constraint on the resulting distributions is that their marginal means must be equal, which results in a nonparametric model that can be difficult to use in applications. Mixtures of Dirichlet distributions have been proposed for use as a semiparametric model, but fitting them is awkward. In this article, we propose a new approach to the use of Dirichlet mixtures, based on tilting, to ensure that the moment conditions are satisfied. We show that these tilted mixtures are dense in the full nonparametric family, are well defined in all dimensions, and allow the probabilistic clustering of extreme events. In order to fit them, we use a fast Markov chain Monte Carlo algorithm that does not require fine-tuning. Its performance is assessed using simulations and an application to financial data.
This article relates the calibration of models to the consistent loss functions for the target functional of the model. Correctly specified models are calibrated. Conversely, we demonstrate that if there is a parameter value that is optimal under all consistent loss functions, then a model is calibrated.
We investigate a B-spline-based estimation procedure for a semiparametric varying coefficient (SPVC) modal regression with an error-prone linear covariate, where the true covariate is unobserved but an ancillary variable is available. The method targets mode values to capture the "most likely" effects rather than mean effects. Varying coefficients are approximated via B-splines, and a deconvolution kernel-based objective function is used for estimation. We establish consistency and asymptotic properties of the estimators under ordinary- and super-smooth error distributions and discuss bandwidth selection. For comparison, asymptotic results for SPVC modal estimators without measurement error are also derived. Efficient implementation relies on a modified fast Fourier transform and a mode expectation-maximization algorithm. Monte Carlo simulations and an empirical application demonstrate the finite-sample performance and practical utility of the proposed method. Cet article propose une m & eacute;thode d'estimation bas & eacute;e sur les B-splines pour un mod & egrave;le de r & eacute;gression modale semi-param & eacute;trique & agrave; coefficients variables (SPVC), lorsqu'une covariable lin & eacute;aire est mesur & eacute;e avec erreur. La vraie covariable n'est pas observable. On dispose & agrave; la place d'une variable auxiliaire. Contrairement aux m & eacute;thodes classiques qui visent la moyenne, cette approche cible le mode de la distribution conditionnelle. Elle capture ainsi les effets les plus probables. Les coefficients variables sont approxim & eacute;s par des B-splines. L'estimation repose sur une fonction objective construite avec un noyau de d & eacute;convolution. Les auteurs prouvent la consistance et les propri & eacute;t & eacute;s asymptotiques des estimateurs. Ces r & eacute;sultats s'appliquent que les erreurs de mesure soient super-lisses ou ordinaires. Ils examinent & eacute;galement le choix de la fen & ecirc;tre de lissage. & Agrave; titre de comparaison, ils & eacute;tablissent aussi les propri & eacute;t & eacute;s asymptotiques des estimateurs modaux SPVC dans le cas o & ugrave; il n'y a pas d'erreur de mesure. Pour l'impl & eacute;mentation, la m & eacute;thode utilise une transform & eacute;e de Fourier rapide modifi & eacute;e et un algorithme d'esp & eacute;rance-maximisation (EM) adapt & eacute; & agrave; l'estimation modale. Des simulations Monte-Carlo et une application sur donn & eacute;es r & eacute;elles montrent les bonnes performances de la m & eacute;thode en & eacute;chantillon fini et son int & eacute;r & ecirc;t pratique.
We develop a model for credit rating migration that accounts for the impact of economic state fluctuations on default probabilities. The joint process for the economic state and the rating is modelled as a time-homogeneous Markov chain. While the rating process itself possesses the Markov property only under restrictive conditions, methods from Markov theory can be used to derive the rating process' asymptotic behaviour. We use the mathematical framework to formalise and analyse different rating philosophies, such as point-in-time (PIT) and through-the-cycle (TTC) ratings. Furthermore, we introduce stochastic orders on the bivariate process' transition matrix to establish a consistent notion of "better" and "worse" ratings. Finally, the construction of PIT and TTC ratings is illustrated on a Merton-type firm-value process.
Granger causal inference is a contentious but widespread method used in fields ranging from economics to neuroscience. The original definition addresses the notion of causality in time series by establishing functional dependence conditional on a specified model. Adaptation of Granger causality to nonlinear data remains challenging, and many methods apply in-sample tests that do not incorporate out-of-sample predictability, leading to concerns of model overfitting. To allow for out-of-sample comparison, a measure of functional connectivity is explicitly defined using permutations of the covariate set. Artificial neural networks serve as featurizers of the data to approximate any arbitrary, nonlinear relationship, and consistent estimation of the variance for each permutation is shown under certain conditions on the featurization process and the model residuals. Performance of the permutation method is compared to penalized variable selection, naive replacement, and omission techniques via simulation, and it is applied to neuronal responses of acoustic stimuli in the auditory cortex of anesthetized rats. Targeted use of the Granger causal framework, when prior knowledge of the causal mechanisms in a dataset are limited, can help to reveal potential predictive relationships between sets of variables that warrant further study.
We analyze the effect of regulatory capital constraints on financial stability in a large homogeneous banking system using a mean-field game (MFG) model. Each bank holds cash and a tradable risky asset. Banks choose absolutely continuous trading rates in order to maximize expected terminal equity, with trades subject to transaction costs. Capital regulation requires equity to exceed a fixed multiple of the position in the tradable asset; breaches trigger forced liquidation. The asset drift depends on changes in average asset holdings across banks, so aggregate de-leveraging creates contagion effects, leading to an MFG. We discuss the coupled forward-backward partial differential equation (PDE) system characterizing equilibria of the MFG and solve the constrained MFG numerically. Experiments demonstrate that capital constraints accelerate de-leveraging and limit risk-bearing capacity. In some regimes, simultaneous breaches trigger liquidation cascades. Finally, we discuss several policy mechanisms for enhancing financial stability. Nous analysons l'effet des contraintes de capital r & eacute;glementaire sur la stabilit & eacute; financi & egrave;re dans un grand syst & egrave;me bancaire & agrave; l'aide d'une formulation en jeu & agrave; champ moyen (JCM). Chaque banque d & eacute;tient de la liquidit & eacute; et un actif risqu & eacute; n & eacute;gociable. Les banques choisissent des vitesses de n & eacute;gociation absolument continues afin de maximiser l'esp & eacute;rance de leur valeur terminale, sous des co & ucirc;ts de transaction. La r & eacute;gulation impose que les fonds propres d & eacute;passent un multiple fix & eacute; de l'actif n & eacute;gociable; toute violation entra & icirc;ne une liquidation forc & eacute;e. La d & eacute;rive de l'actif d & eacute;pend des positions moyenness, de sorte que le d & eacute;sendettement agr & eacute;g & eacute; engendre des effets de contagion, menant & agrave; un JCM. Nous discutons le syst & egrave;me d'EDP coupl & eacute;es caract & eacute;risant les & eacute;quilibres du JCM, et nous r & eacute;solvons le JCM contraint num & eacute;riquement. Les exp & eacute;riences montrent que les contraintes acc & eacute;l & egrave;rent le d & eacute;sendettement et r & eacute;duisent la capacit & eacute; de portage du risque. Dans certains r & eacute;gimes, des violations simultan & eacute;es d & eacute;clenchent des cascades de liquidation. Enfin, nous & eacute;tudions plusieurs m & eacute;canismes pour renforcer la stabilit & eacute; financi & egrave;re.
In extreme value theory, the presence of asymptotic independence signifies that joint extreme events across multiple variables are unlikely. Although well understood in a bivariate context, the concept remains relatively unexplored when addressing the nuances of simultaneous occurrence of extremes in higher dimensions. In this article, we propose a notion of mutual asymptotic independence to capture the behaviour of joint extremes in dimensions larger than two and contrast it with the classical notion of (pairwise) asymptotic independence. Additionally, we define -wise asymptotic independence, which captures the tail dependence in between pairwise and mutual asymptotic independence. The concepts are compared using examples of Archimedean, Gaussian, and Marshall-Olkin copulas, among others. Finally, we discuss the implications of these new notions of asymptotic independence on assessing the risk in complex systems under distributional ambiguity. Dans le cadre de la th & eacute;orie des valeurs extr & ecirc;mes, la notion d'ind & eacute;pendance asymptotique traduit le fait que la survenue simultan & eacute;e d'& eacute;v & eacute;nements extr & ecirc;mes sur plusieurs variables est peu probable. Bien que ce concept soit bien maitris & eacute; en dimension bivari & eacute;e, son & eacute;tude demeure limit & eacute;e lorsqu'il s'agit d'appr & eacute;hender la cooccurrence d'& eacute;v & eacute;nements extr & ecirc;mes dans des dimensions sup & eacute;rieures. Cet article introduit la notion d'ind & eacute;pendance asymptotique mutuelle pour ca-ract & eacute;riser le comportement conjoint des valeurs extr & ecirc;mes en dimension sup & eacute;rieure & agrave; deux, par opposition & agrave; la notion classique d'ind & eacute;pendance asymptotique par paires. Les auteurs introduisent & eacute;galement la notion d'ind & eacute;pendance asymptotique entre k-uplets, qui permet de mettre en & eacute;vidence des formes interm & eacute;diaires de d & eacute;pendance extr & ecirc;me, situ & eacute;es entre l'ind & eacute;pendance asymptotique par paires et l'ind & eacute;pendance asymptotique mutuelle. Ces diff & eacute;rentes notions sont illustr & eacute;es et compar & eacute;es & agrave; travers plusieurs exemples, notamment les copules archim & eacute;diennes, gaussiennes et de Marshall-Olkin. Enfin, les implications de ces nouvelles notions d'ind & eacute;pendance asymptotique sont examin & eacute;es dans le contexte de l'& eacute;valuation du risque pour des syst & egrave;mes complexes affect & eacute;s par une incertitude sur la loi de probabilit & eacute; sous-jacente.
We establish the consistency and the asymptotic distribution of the least squares estimators of the coefficients of a subset vector autoregressive process with exogenous variables (VARX). Using a martingale central limit theorem, we derive the asymptotic normal distribution of the estimators. Diagnostic checking is discussed using kernel-based spectral density estimators. Multi-step conditional prediction intervals are developed. In a small empirical study, subset VARX models are fitted and compared when lags are chosen according to some popular model selection techniques. Coverage properties of the prediction intervals are illustrated. An application to the monthly production of cheese in Canada for the time period 2003-2021 illustrates the methodology. Nous & eacute;tablissons la convergence en probabilit & eacute; et la distribution asymptotique des estimateurs des moindres carr & eacute;s des coefficients d'un processus autor & eacute;gressif vectoriel avec variables exog & egrave;nes (VARX) et avec s & eacute;lection d'indices. Utilisant un th & eacute;or & egrave;me central limite pour martingales, nous obtenons la distribution asymptotique normale des estimateurs. La qualit & eacute; d'ajustement est discut & eacute;e, utilisant des estimateurs de la densit & eacute; spectrale bas & eacute;s sur la m & eacute;thode du noyau. Des intervalles de pr & eacute;vision conditionnels & agrave; horizons multiples sont d & eacute;velopp & eacute;s. Dans une br & egrave;ve & eacute;tude empirique, des mod & egrave;les VARX avec s & eacute;lection d'indices sont ajust & eacute;s et compar & eacute;s lorsque les d & eacute;lais sont choisis selon des m & eacute;thodes de s & eacute;lection de mod & egrave;les populaires. Les propri & eacute;t & eacute;s de couverture des intervalles de pr & eacute;vision sont illustr & eacute;es. Une application & agrave; la production mensuelle de fromage au Canada pour la p & eacute;riode 2003-2021 illustre la m & eacute;thodologie.
We investigate the family of cross-classified sampling designs across an arbitrary number of dimensions. We introduce a variance decomposition that enables the derivation of general asymptotic properties for these designs and the development of straightforward and asymptotically unbiased variance estimators. Additionally, we demonstrate the suitability of weighted bootstrap techniques for cross-classified sampling, given the availability of a weighted bootstrap technique in each dimension. Our conclusions are supported by an extensive simulation study. Finally, we apply the proposed methods to a French longitudinal survey conducted on children. Nous & eacute;tudions les plans produits dans un nombre arbitraire de dimensions. Nous introduisons une d & eacute;composition de la variance permettant d'& eacute;tablir des propri & eacute;t & eacute;s asymptotiques g & eacute;n & eacute;rales de ces plans et de d & eacute;velopper des estimateurs de variance simples et asymptotiquement sans biais. De plus, nous montrons qu'il est possible de produire un estimateur bootstrap & agrave; poids pour les plans produits, pourvu que l'on dispose d'une m & eacute;thode bootstrap & agrave; poids dans chaque dimension. Une & eacute;tude par simulation permet de v & eacute;rifier nos r & eacute;sultats. Enfin, nous appliquons & eacute;galement les m & eacute;thodes propos & eacute;es & agrave; l'Etude Longitudinale Fran & ccedil;aise depuis l'Enfance.
A methodology for modelling and forecasting univariate time series using non‐Gaussian ARMA and seasonal ARIMA models based on D‐vine copulas is proposed. By combining a parametric D‐vine process to describe serial dependence with a nonparametric or parametric model of the marginal distribution, the method offers improved modelling and distributional forecasting for time series that have a non‐Gaussian distribution and a nonlinear dependence on past values. While D‐vine copula‐based models of univariate time series are known to generalize the classical Gaussian autoregressive (AR) model, an innovative method of parametrization based on the Kendall partial autocorrelation function is shown to permit models that generalize any ARMA model. Simulations and examples of real data show the forecasting advantages of using non‐Gaussian and nonlinear serial dependence structures, as well as the advantages of improved marginal modelling that are offered by a copula approach.
Motivated by a recent trial using MRI scans to compare treatments for nasopharyngeal carcinoma, we recognize that imaging predictors can significantly improve treatment selection. However, such models are built retrospectively under equal randomization, not integrated in real time to inform adaptive treatment allocation. This article proposes a covariate-adjusted response-adaptive (CARA) randomization framework that prospectively incorporates imaging data. Using supervised functional principal component analysis (sFPCA), we extract features from patient images to enable adaptive randomization based on imaging covariates. By dynamically adjusting randomization probabilities to favour more effective treatments as data accumulate, our method aims to enhance patient outcomes within the trial, supporting an ethical trial design. Simulations show that CARA with imaging covariates allocates more patients to better treatments while maintaining good statistical power and accuracy.