
This paper proposes a new approach to fit a linear regression for symbolic internal-valued variables, which improves both the Center Method suggested by Billard and Diday in and the Center and Range Method suggested by Lima-Neto, E.A. and De Carvalho, F.A.T. in . Just in the Centers Method and the Center and Range Method, the new methods proposed fit the linear regression model on the midpoints and in the half of the length of the intervals as an additional variable (ranges) assumed by the predictor variables in the training data set, but to make these fitments in the regression models, the methods Ridge Regression, Lasso, and Elastic Net proposed by Tibshirani, R. Hastie, T., and Zou H in are used. The prediction of the lower and upper of the interval response (dependent) variable is carried out from their midpoints and ranges, which are estimated from the linear regression models with shrinkage generated in the midpoints and the ranges of the interval-valued predictors. Methods presented in this document are applied to three real data sets cardiologic interval data set, Prostate interval data set and US Murder interval data set to then compare their performance and facility of interpretation regarding the Center Method and the Center and Range Method. For this evaluation, the root-mean-squared error and the correlation coefficient are used. Besides, the reader may use all the methods presented herein and verify the results using the RSDA package written in R language, that can be downloaded and installed directly from CRAN .
Résumé. Dans cet article, nous comparons différentes méthodes qui permettent de trouver dans un ensemble fini S, sur lequel on a défini une dissimilarité, un sous-ensemble d’éléments qui s’apparient au mieux aux éléments d’un ensemble cible C également muni d’une dissimilarité. La qualité de l’appariement est mesurée par la distance observée entre les matrices de dissimilarités définies sur les sous-ensembles de S et sur C. Différents algorithmes sont présentés et les performances sont commentées.
The use of structural equation models in management science, especially in marketing, is a methodological and empirical promising axis and innovative direction toward development of the theory, based on a set of approaches and advanced techniques. Therefore, this article mainly focuses on explaining the value and interest of these second generation methods in the validation of measures and causality models, and the specification of the theoretical constructs and relationships studied simultaneously. After presenting an overview on the conceptual basis and procedure of carrying out a structural equation model, the second part of this article attempts to expose the common practice of the methods adopted by researchers in marketing. Empirically, it seems important to propose concrete and illustrative example dealing with the study of the relationship among customers’ service quality, satisfaction and loyalty to their telephone service providers. Finally, an investigation was made to 223 respondents in order to validate a causal model in services field.
The problem of the proper dimension of a Multiple Correspondence Analysis (MCA) is discussed, based on both the re-evaluation of the explained inertia sensu Benzecri (1979) and Greenacre (2006) and a test proposed by Ben Ammou and Saporta (1998). This leads to the consideration of a better reconstruction of the off-diagonal sub-tables of the Burt’s table crossing the nominal characters taken into the account. Thus, Greenacre (1988) Joint Correspondence Analysis (JCA) is introduced and the results obtained on two applications are shown and the quality of reconstruction of both MCA and JCA solutions are compared to the Simple Correspondence Analysis results of the two-way tables. It results that JCA’s reduced-dimensional reconstruction is much better than the MCA’s one, that reveals highly biased and non-monotonous.
La régression sur composantes principales (RCP) est une régression sur les facteurs d’une ACP préalablement effectuée sur des variables initi alement corrélées. L’utilisation de l’ACP permet de remplacer les variables initiales, par de s composantes principales qui conservent la quasi-totalité de l’information, et qui présentent l’avantage d’être non corrélées. Ces composantes, sont prises comme variables explicativ es pour une régression linéaire multiple. La qualité de la modélisation par RCP reste affecté e par l’existence de bruit dans les variables initiales. Nous proposons dans ce travail un débrui tage des données par ondelettes (wavelets) permettant de séparer le signal du bruit sans perte d’information. Nous montrons, sur des données boursières françaises, que l’élimination du bruit sur les composantes principales par un seuillage doux à base d’ondelettes améliore la q u lité d’ajustement du modèle de régression ainsi que la qualité des prévisions. Nou s c nfirmons le résultat par simulation.
Partial Least Squares regression and Principal Comp onents Regression make possible to relate a set of dependant variables Y to a set of i ndependent variables X, when there is multicollinearity. This paper suggests a new approa ch for analyzing the net interest margin. After usi ng the PCR method, the determinants of the net interes t margin have been viewed through a PLS model.
Based on the theoretical structure of hierarchical c l ssification to build the tree or dendrogram, is s hown the theoretical relationship of geometrical hierarchica l distances for a sequence of partial hierarchies w here two partial and equal hierarchies exist in the election of clas ses to be added, then the partial hierarchy to be a dd d depends on geometric distances shown by partial hierarchies re garding the third class. Theoretical development is exemplified through applications with data from the effect of a tmospheric corrosion of structural steel in civil i nfrastructure in Mexico City and the assessment of teaching performan ce for postgraduate studies in Mexico.
Resume Cet article est une introduction au domaine de la phylogenie moleculaire et en particulier a la robustesse des arbres phylogenetiques. Nous commencons par une breve presentation historique du domaine avant de passer en revue les methodes de reconstruction les plus populaires. Nous nous interessons tout particulierement a la methode du maximum de vraisemblance. Cette methode necessite de construire un modele probabiliste d’evolution des macromolecules biologiques mais fournit en contrepartie un cadre statistique propice a quantifier la variabilite de l’arbre estime. Nous presentons tout d’abord les modeles d’evolution couramment utilise, puis le calcul de la vraisemblance avant de montrer que la nature discrete de l’arbre rend caducs les outils traditionnels d’etude de la variabilite.