Changes in health-related quality of life (HRQoL) over time are not necessarily homogeneous within a population of interest. Our study aim was twofold: to determine homogeneous patient subpopulations distinguished by HRQoL trajectories, and to identify the particular patient profile associated with each subpopulation. To classify patients according to HRQoL dimension scores, we compared mixtures of linear mixed models (LMMs) classically applied to scores defined by the EORTC procedure, and mixtures of random effect cumulative models (CMs) applied to scores treated as ordinal variables. A simulation study showed that the mixture of LMMs overestimated the number of subpopulations and was less able to correctly classify patients than the mixture of CMs. Considering HRQoL scores as ordinal rather than continuous variables is relevant when classifying patients. The mixture of CMs for ordinal scores is able to identify homogeneous subpopulations and their associated trajectories. The application focused on changes over time in HRQoL data (collected using the EORTC QLQ-C30 questionnaire) from 132 breast cancer patients from the Moral study. Once the classification is obtained only from HRQoL scores, class membership was then explained through a logistic regression model, given a large panel of variables collected at baseline. Analysis of data revealed that deterioration over time of role functioning and insomnia was closely related to patient anxiety: anxiety at baseline is a prognostic factor for a poor level and/or a deterioration over time of HRQoL. For functional dimensions, large tumor size and high education level were associated with worse HRQoL scores.
Click-Through-Rate (CTR) prediction is one of the most important challenges in the advertisement field. Nevertheless, it is essential to understand beforehand the voluminous and heterogeneous data structure. Here we introduce a novel CTR prediction method using a mixture of generalized linear models (GLMs). First, we develop a model-based clustering method dedicated to publicity campaign time-series, i.e. non-Gaussian longitudinal data. Secondely, we consider two CTR predictive models derived from the inferred clustering. The clustering step improves the CTR prediction performance both on simulated and real data. An R package binomialMix for mixture of binomial and longitudinal data is available on CRAN.
Today, social media is increasingly used by patients to openly discuss their health. Mining automatically such data is a challenging task because of the non-structured nature of the text and the use of many abbreviations and the slang terms. Our goal is to use Patient Authored Text to build a French Consumer Health Vocabulary on breast cancer field, by collecting various kinds of non-experts' expressions that are related to their diseases and then compare them to biomedical terms used by health care professionals. We combine several methods of the literature based on linguistic and statistical approaches to extract candidate terms used by non-experts and to link them to expert terms. We use messages extracted from the forum on cancerdusein org and a vocabulary dedicated to breast cancer elaborated by the Institut National Du Cancer. We have built an efficient vocabulary composed of 192 validated relationships and formalized in Simple Knowledge Organization System ontology.
Health‐related quality of life (HRQoL) data are measured via patient questionnaires, completed by the patients themselves at different time points. We focused on oncology data gathered through the use of European Organization for Research and Treatment of Cancer questionnaires, which decompose HRQoL into several functional dimensions, several symptomatic dimensions, and the global health status (GHS). We aimed to perform a global analysis of HRQoL and reduce the number of analyses required by using a two‐step approach. First, a structural equation model (SEM) was used for each time point; in these models, the GHS is explained by two latent variables. Each latent variable is a factor that summarizes, respectively, the functional dimensions and the symptomatic dimensions to the global measurement. This is achieved through the maximization of the likelihood of each SEM using the EM algorithm, which has the advantage of giving an estimation of the subject‐specific factors and the influence of additional explanatory variables. Then, to consider the longitudinal aspect, the GHS variable and the two factors were concatenated for each patient visit at which the questionnaire was completed. The GHS and the two factors estimated in the first step can then be explained by additional explanatory variables using a linear mixed model.
BACKGROUND:Social media dedicated to health are increasingly used by patients and health professionals. They are rich textual resources with content generated through free exchange between patients. We are proposing a method to tackle the problem of retrieving clinically relevant information from such social media in order to analyze the quality of life of patients with breast cancer.OBJECTIVE:Our aim was to detect the different topics discussed by patients on social media and to relate them to functional and symptomatic dimensions assessed in the internationally standardized self-administered questionnaires used in cancer clinical trials (European Organization for Research and Treatment of Cancer [EORTC] Quality of Life Questionnaire Core 30 [QLQ-C30] and breast cancer module [QLQ-BR23]).METHODS:First, we applied a classic text mining technique, latent Dirichlet allocation (LDA), to detect the different topics discussed on social media dealing with breast cancer. We applied the LDA model to 2 datasets composed of messages extracted from public Facebook groups and from a public health forum (cancerdusein.org, a French breast cancer forum) with relevant preprocessing. Second, we applied a customized Jaccard coefficient to automatically compute similarity distance between the topics detected with LDA and the questions in the self-administered questionnaires used to study quality of life.RESULTS:Among the 23 topics present in the self-administered questionnaires, 22 matched with the topics discussed by patients on social media. Interestingly, these topics corresponded to 95% (22/23) of the forum and 86% (20/23) of the Facebook group topics. These figures underline that topics related to quality of life are an important concern for patients. However, 5 social media topics had no corresponding topic in the questionnaires, which do not cover all of the patients' concerns. Of these 5 topics, 2 could potentially be used in the questionnaires, and these 2 topics corresponded to a total of 3.10% (523/16,868) of topics in the cancerdusein.org corpus and 4.30% (3014/70,092) of the Facebook corpus.CONCLUSIONS:We found a good correspondence between detected topics on social media and topics covered by the self-administered questionnaires, which substantiates the sound construction of such questionnaires. We detected new emerging topics from social media that can be used to complete current self-administered questionnaires. Moreover, we confirmed that social media mining is an important source of information for complementary analysis of quality of life.
Background The use of health-related quality of life (HRQoL) as an endpoint in cancer clinical trials is growing rapidly. Hence, research into the statistical approaches used to analyze HRQoL data is of major importance, and could lead to a better understanding of the impact of treatments on the everyday life and care of patients. Amongst the models that are used for the longitudinal analysis of HRQoL, we focused on the mixed models from item response theory, to directly analyze raw data from questionnaires. Methods We reviewed the different item response models for ordinal responses, using a recent classification of generalized linear models for categorical data. Based on methodological and practical arguments, we then proposed a conceptual selection of these models for the longitudinal analysis of HRQoL in cancer clinical trials. Results To complete comparison studies already present in the literature, we performed a simulation study based on random part of the mixed models, so to compare the linear mixed model classically used to the selected item response models. As expected, the sensitivity of the item response models to detect random effects with lower variance is better than that of the linear mixed model. We then used a cumulative item response model to perform a longitudinal analysis of HRQoL data from a cancer clinical trial. Conclusions Adjacent and cumulative item response models seem particularly suitable for HRQoL analysis. In the specific context of cancer clinical trials and the comparison between two groups of HRQoL data over time, the cumulative model seems to be the most suitable, given that it is able to generate a more complete set of results and gives an intuitive illustration of the data.
En oncologie, la qualité de vie relative à la santé (QdV) fait l’objet de nombreuses analyses afin d’améliorer la prise en charge du patient. Dans ce travail, nous nous intéressons à l’étude des trajectoires de QdV et à leur hétérogénéité au sein d’une même cohorte de patient. L’objectif est, d’une part, de déterminer des sous-populations homogènes non observées (classes latentes) au travers des trajectoires de QdV des patients, et d’autre part, de déterminer si des variables socio-démographiques et psychologiques peuvent expliquer l’appartenance aux différentes classes latentes dans le but d’identifier des profils types de patient. Cette approche est illustrée sur les données de l’étude Moral concernant 132 patientes atteintes d’un cancer du sein. Les mesures de QdV sont issues de la collecte de l’auto-questionnaire EORTC QLQ-C30 à différentes visites au cours du traitement et du suivi. Ce questionnaire décompose la QdV en dimensions fonctionnelles et symptomatiques et en un statut global de santé. Le recueil des données incluant des variables socio-démographiques (âge, niveau d’études, situation familiale…), médicales (taille de la tumeur, grade…), ainsi que des données psychosociales (anxiété, dépression, soutien social, optimisme…) est réalisé à six visites (à la chirurgie puis à 1, 4, 7, 10 et 13 mois). La classification s’effectue par un mélange de modèles mixtes, où la distinction entre chaque composante est réalisée suivant la trajectoire de QdV. Pour les dimensions de QdV ne comportant qu’un seul item, un mélange de modèles issus de la théorie de réponse à l’item est utilisé. Pour les autres dimensions, nous utilisons un mélange de modèles linéaires mixtes. Ces mélanges permettent pour chaque dimension considérée d’obtenir une partition de la population d’intérêt ainsi que la trajectoire moyenne associée à chaque sous-population. Un modèle logistique multinomiale a été utilisé pour expliquer l’appartenance aux différentes classes latentes et ainsi proposer un profil type pouvant expliquer ou prédire l’évolution de la QdV. Concernant les dimensions de QdV « insomnie » et « perte d’appétit », l’état anxieux de la patiente à la chirurgie (baseline) est prédictif de l’appartenance à la classe latente ayant la trajectoire de QdV la moins favorable. Les patientes plus jeunes ont une fonction émotionnelle plus faible à la baseline (p = 0,001) que les patientes plus âgées. Les patientes avec un niveau d’étude supérieur ou égal au baccalauréat sont plus exposées à une dégradation de leur capacité fonctionnelle (p = 0,003) et de leur niveau de QdV global (p ≤ 0,001). L’approche utilisée permet de faire émerger des profils particuliers de patients définis par des trajectoires latentes de QdV. Ces profils peuvent ensuite être caractérisés par des variables explicatives permettant de prédire l’appartenance à une classe latente afin de mieux prendre en charge le patient par la suite. L’analyse des données de l’étude Moral met en évidence certaines caractéristiques socio-démographiques et psychologiques des patientes à l’inclusion qui peuvent prédire l’évolution de leur QdV.
OPA1 mutations are responsible for autosomal dominant optic atrophy (ADOA), a progressive blinding disease characterized by retinal ganglion cell (RGC) degeneration and large phenotypic variations, the underlying mechanisms of which are poorly understood. OPA1 encodes a mitochondrial protein with essential biological functions, its main roles residing in the control of mitochondrial membrane dynamics as a pro-fusion protein and prevention of apoptosis. Considering recent findings showing the importance of the mitochondrial fusion process and the involvement of OPA1 in controlling steroidogenesis, we tested the hypothesis of deregulated steroid production in retina due to a disease-causing OPA1 mutation and its contribution to the visual phenotypic variations. Using the mouse model carrying the human recurrent OPA1 mutation, we disclosed that Opa1 haploinsufficiency leads to very high circulating levels of steroid precursor pregnenolone in females, causing an early-onset vision loss, abolished by ovariectomy. In addition, steroid production in retina is also increased which, in conjunction with high circulating levels, impairs estrogen receptor expression and mitochondrial respiratory complex IV activity, promoting RGC apoptosis in females. We further demonstrate the involvement of Muller glial cells as increased pregnenolone production in female cells is noxious and compromises their role in supporting RGC survival. In parallel, we analyzed ophthalmological data of a multicentre OPA1 patient cohort and found that women undergo more severe visual loss at adolescence and greater progressive thinning of the retinal nerve fibres than males. Thus, we disclosed a gender-dependent effect on ADOA severity, involving for the first time steroids and Müller glial cells, responsible for RGC degeneration.
En oncologie, la qualité de vie relativea la santé (QdV) est un critere secondaire dans les essais cliniques mais son l’analyse longitudinale reste complexe. La QdV est mesuréea travers des questionnaires que remplissent les patientsa différentes visites du traitement et au cours du suivi. La structure particuliere du questionnaire QLQ-C30 de l’EORTC décompose la QdV en un groupe de dimensions fonctionnelles, un groupe de dimensions symptomatiques et le” Statut global de santé”(GHS, Global Health Status). Par un Modelea Équation Structurelle (SEM, Structural Equation Model), l’objectif est d’expliquer le GHS par les autres dimensions. Cette modélisationa équation structurelle est réaliséea chaque visite ou la variable GHS est expliquée par deux variables latentes. Chaque variable latente résume respectivement le groupe de variables fonctionnelles et le groupe de variables symptomatiques. Puis une maximisation de la vraisemblance par algorithme EM de chacun des modeles, permet d’obtenira chaque visite une estimation des facteurs. L’analyse longitudinale est alors réalisée par l’intermédiaire d’un modele linéaire mixte sur la concaténation de la variable GHS et des facteurs estimésa chaque visite. Cette modélisation permet de prendre en compte la variabilité intra-individuelle avec les effets aléatoires et l’influence d’éventuelles covariables tel que le traitement. Nous présenterons une application de cette approche sur des données réelles issues d’un essai clinique en cancérologie.
De nos jours, les medias sociaux sont de plus en plus utilises par les patients et les professionnels de sante. Les patients, generalement profanes dans le domaine medical, utilisent de l'argot, des abreviations et un vocabulaire qui leur est propre lors de leurs echanges. Pour analyser automatiquement les textes des reseaux sociaux, l'acquisition de ce vocabulaire spe-cifique est necessaire. En nous appuyant sur un corpus de documents issus de messages de medias sociaux de type forums ou Facebook, nous decrivons la construction d'une ressource lexicale qui aligne le vocabulaire des patients a celui des professionnels de sante. Nous utili-sons plusieurs methodes prenant en compte les aspects linguistiques et statistiques proposees dans la litterature pour construire cette ressource et nous la transformons en une ontologie SKOS (Simple Knowledge Organization System). Ce travail permettra, d'une part d'ameliorer la recherche d'informations dans les forums de sante et d'autre part, de faciliter l'elaboration d'etudes statistiques basees sur les informations extraites de ces forums.
De nos jours, les medias sociaux sont de plus en plus utilises par les patients et les professionnels de sante. Il s'agit d'une ressource textuelle riche, generee par les tres nombreux echanges entre patients et, dans certains cas, professionnels de sante. Dans cet article, nous utilisons le modele d'apprentissage non supervise connu sous le nom de LDA (Allocation de Dirichlet Latente) afin de detecter les differents themes abordes sur les forums de sante et les reseaux sociaux par les patients. Notre objectif est de reperer les nouveaux themes directement issus des preoccupations des patientes atteintes de cancer du sein et de les comparer aux themes existant dans les auto-questionnaires proposes dans les essais cliniques en oncologie. Mots-
In this work, we propose a new estimation method of a Structural Equation Model. Our method is based on the EM likelihood-maximization algorithm. We show that this method provides estimators, not only of the coefficients of the model, but also of its latent factors. Through a simulation study, we investigate how fast and accurate the method is, and then apply it to real environmental data.
Introduction Dans les essais cliniques en cancerologie, la qualite de vie relative a la sante (QdV) est un critere essentiel pour evaluer l’efficacite d’une prise en charge. Cependant, son analyse longitudinale est non standardisee et un des freins conceptuels est son aspect multidimensionnel. L’objectif de ce travail est de proposer une nouvelle approche qui permet de tenir compte conjointement de l'aspect longitudinal et de la nature multidimensionnelle de la QdV. Methodes Les mesures de QdV sont realisees par l’intermediaire d’auto-questionnaires collectes a differentes visites du traitement et du suivi. Le questionnaire standard en Europe (l’EORTC-QLQ-C30) decompose la QdV en 15 dimensions : 5 dimensions fonctionnelles, 9 symptomatiques et le statut global de sante (GHS). A chaque temps, un modele a equations structurelles (SEM) est construit ou le GHS (unidimensionnel) est explique par deux variables latentes resumant le statut fonctionnel et le statut symptomatique, respectivement reflete par les dimensions fonctionnelles et par les dimensions symptomatiques. Pour tenir compte de l'aspect longitudinal, la variable GHS et les deux facteurs seront concatenes sur l'ensemble des visites. Au travers d’un modele lineaire mixte (LMM), la GHS est expliquee par des variables explicatives (dont les deux variables latentes) et un effet individu (effet aleatoire). Ceci est rendu possible par la maximisation de la vraisemblance de chaque SEM utilisant l'algorithme EM. Ce dernier donne une estimation des facteurs ce qui permet de les reintroduire dans le LMM. Resultats et conclusion Cette methode est developpee sous le logiciel R et est illustree via les donnees issues de l’etude CO-HO-RT. Cet essai clinique de phase II randomise est mene sur 150 patientes et evalue les sequences de traitement de radiotherapie (concomitante ou sequentielle) associee a une hormonotherapie. Notre approche est proposee pour eviter la multiplicite des tests induite par l’analyse independante de chaque dimension.
Introduction. A new longitudinal statistical approach was compared to the classical methods currently used to analyze health-related quality-of-life (HRQoL) data. The comparison was made using data in patients with metastatic pancreatic cancer. Methods. Three hundred forty-two patients from the PRODIGE4/ACCORD 11 study were randomly assigned to FOLFIRINOX versus gemcitabine regimens. HRQoL was evaluated using the European Organization for Research and Treatment of Cancer (EORTC) QLQ-C30. The classical analysis uses a linear mixed model (LMM), considering an HRQoL score as a good representation of the true value of the HRQoL, following EORTC recommendations. In contrast, built on the item response theory (IRT), our approach considered HRQoL as a latent variable directly estimated from the raw data. For polytomous items, we extended the partial credit model to a longitudinal analysis (longitudinal partial credit model [LPCM]), thereby modeling the latent trait as a function of time and other covariates. Results. Both models gave the same conclusions on 11 of 15 HRQoL dimensions. HRQoL evolution was similar between the 2 treatment arms, except for the symptoms of pain. Indeed, regarding the LPCM, pain perception was significantly less important in the FOLFIRINOX arm than in the gemcitabine arm. For most of the scales, HRQoL changes over time, and no difference was found between treatments in terms of HRQoL. Discussion. The use of LMM to study the HRQoL score does not seem appropriate. It is an easy-to-use model, but the basic statistical assumptions do not check. Our IRT model may be more complex but shows the same qualities and gives similar results. It has the additional advantage of being more precise and suitable because of its direct use of raw data.
Introduction Dans les essais cliniques en cancérologie, la qualité de vie relative à la santé (QdV) est un critère essentiel pour évaluer l’efficacité d’une prise en charge. Cependant, son analyse longitudinale est non standardisée et un des freins conceptuels est son aspect multidimensionnel. L’objectif de ce travail est de proposer une nouvelle approche qui permet de tenir compte conjointement de l'aspect longitudinal et de la nature multidimensionnelle de la QdV. Méthodes Les mesures de QdV sont réalisées par l’intermédiaire d’auto-questionnaires collectés à différentes visites du traitement et du suivi. Le questionnaire standard en Europe (l’EORTC-QLQ-C30) décompose la QdV en 15 dimensions : 5 dimensions fonctionnelles, 9 symptomatiques et le statut global de santé (GHS). A chaque temps, un modèle à équations structurelles (SEM) est construit où le GHS (unidimensionnel) est expliqué par deux variables latentes résumant le statut fonctionnel et le statut symptomatique, respectivement reflété par les dimensions fonctionnelles et par les dimensions symptomatiques. Pour tenir compte de l'aspect longitudinal, la variable GHS et les deux facteurs seront concaténés sur l'ensemble des visites. Au travers d’un modèle linéaire mixte (LMM), la GHS est expliquée par des variables explicatives (dont les deux variables latentes) et un effet individu (effet aléatoire). Ceci est rendu possible par la maximisation de la vraisemblance de chaque SEM utilisant l'algorithme EM. Ce dernier donne une estimation des facteurs ce qui permet de les réintroduire dans le LMM. Résultats et conclusion Cette méthode est développée sous le logiciel R et est …
This paper describes the methods we submitted to the DEFT 2015 Challenge (Text Mining Challenge). This eleventh edition concerned the analysis of opinions, sentiments and emotions expressed in French tweets. Three tasks have been proposed, we participated to task 1 which concerned the classification of tweets according to their polarities, to task 2.1 concerning the identification of the generic class of information expressed in the tweet, and finally to task 2.2 that concerned the identification of the specific class of opinion, sentiment or emotion. We proposed supervised methods based on support vector machines (SVM) using several types of attributes such as word n-grams, character n-grams, most common syntactic patterns, etc. Moreover, we constructed and used two French lexicons of sentiments and emotions.
Pascal Poncelet合作论文数University Montpellier 2 - LIRMM2