Clinical and epidemiological researchers may want to identify or determine groups with similar characteristics that exist within a dataset or population. For this purpose, clustering techniques can be applied. However, clustering can be complex, and results can be misleading if the methods are not applied with fidelity.The aim of this paper is to provide practical guidance to (clinical) researchers in rheumatology and other fields who wish to explore or confirm (sub) groups within their study population, and are considering to apply clustering techniques to (patient) data. We discuss when clustering is useful for addressing your research question (and when it is not), the need to define the cluster concept, the choice of distance measure between 2 observations, the preprocessing of the data, the assessment of clustering tendency, the types of clustering methods, the determination of the number of clusters and end with the evaluation of the derived clustering solutions using visualizations and statistical measures.The following methods are presented in this paper: K-means, partitioning around medoids, hierarchical clustering, Density-Based Spatial Clustering of Applications with Noise, spectral clustering, fuzzy clustering, and latent class analysis. To illuminate the different considerations involved in clustering, we use a case example: secondary data from 895 people with rheumatoid arthritis, spondyloarthritis, and gout who completed the Health Literacy Questionnaire (HLQ). In this case example, we compute, appraise and evaluate clustering solutions using different methods in 6 steps, including statistical measures and expert opinion.The 6 steps proposed in this paper describe a systematic approach to cluster analysis to ensure all important aspects are considered. In the case example, different clustering methods led to different clustering solutions and insights. Qualitative interpretation by content experts was insightful here. Where possible, researchers should apply more than one clustering method fitting to the objective and reflect on eventual differences in solutions.
Nanomaterials can be found in many applications, from daily products to the healthcare industry. Since human can be exposed to nanomaterials through many ways, it is necessary to study the nanomaterials, especially their potential adverse effects on humans. This research was conducted under the European’s Union H2020 NanoInformaTIX project and focused on the dose-response analysis of nanomaterials’ toxicity. This research, focusing on the data of in vitro studies, aimed to model the relationship between the amount of administered nanomaterial and the possible toxic response on the cells using nonlinear models for the dose-response data. The data used as an example consisted of 65 data of nanomaterials which was differentiated by the cell types, on which the Likelihood ratio test was first applied to identify significant monotone trend. Dose-response model fitting was then conducted on the 14 data subsets with significant monotone trends. Several nonlinear models such as the Three-, Four-, and Five-parameter Log-logistic model, Weibull model, and Gompertz model were fitted on the data. As an illustration, the analysis of NM-110 (zinc oxide, uncoated) and NM-102 (titanium dioxide, anatase) was presented. For the NM-102 (titanium dioxide, anatase), the best model was the Weibull model according to the value of the AIC, with the value of ED50 equals 22.710 (95
AIMS:Identifying biomarkers that reflect the complex relationship between the microbiome and health outcomes in microbiome studies is essential for advancing the understanding and improving disease management. While past research was focused on a single biomarker modeling approach, this study extends that work by combining multiple taxa to identify a subset of multiple biomarkers relevant to clinical outcomes. METHODS AND RESULTS:We extend the information theory framework for surrogate endpoint evaluation by applying LASSO and Elastic Net models to identify combinations of taxa as biomarkers for clinical outcomes. Feature selection for the biomarker's construction is done in order to maximize the goodness of fit of the predictive biomarker model. Monte Carlo cross validation is used to enhance the reliability of feature selection. The high salt diet study on mice is used to illustrate the methodology for continuous outcome (tumor size). The top 5 selected genera yielded a correlation of 0.9274 between predicted and observed tumor size, with a 67.92% reduction in uncertainty when the multiple microbiome biomarkers score is known. To illustrate the methodology for binary outcome, the CERTIFI study on Crohn's disease patients treated with ustekinumab is used. A multiple microbiome biomarkers score, constructed using the top 5 selected families, significantly improved prediction of remission 6 weeks after induction treatment (the clinical outcome of interest). CONCLUSIONS:This study presents a unified approach for identifying multiple microbiome biomarkers using penalized regression for clinical outcome prediction. The proposed methods are applied to both continuous and binary outcomes. The method enhances the detection of meaningful biomarkers with potential for personalized treatment and disease management.
Motivation:Insights from integrative multi-omics analyses have fueled demand for innovative computational methods and tools in multi-omics research. However, the scarcity of multi-omics datasets with user-defined signal structures hinders the evaluation of these newly developed tools. SUMO (SimUlating Multi-Omics), an open-source R package, was developed to address this gap by enabling the generation of high-quality factor analysis-based datasets with full control over the dataset's structure such as latent structures, noise, and complexity. Users can configure datasets with distinct and/or shared non-overlapping latent factors, enabling flexible and precise control over the signal structures. Consequently, SUMO allows reproducible testing and validation of methods, fostering methodological innovation. Availability and implementation:The SUMO R package is freely available and accessible on the Comprehensive R Archive Network https://doi.org/10.32614/CRAN.package.SUMO and on GitHub https://github.com/lucp12891/SUMO.git under CC-BY 4.0 license.
Nanomaterials are increasingly used in many applications due to their enhanced properties. To ensure their safety for humans and the environment, nanomaterials need to be evaluated for their potential risk. The risk assessment analysis on the nanomaterials based on animal or in vivo studies is accompanied by several concerns, including animal welfare, time and cost needed for the studies. Therefore, incorporating in vitro studies in the risk assessment process is increasingly considered. To be able to analyze the potential risk of nanomaterial to human health, there are factors to take into account. Utilizing in vitro data in the risk assessment analysis requires methods that can be used to translate in vitro data to predict in vivo phenomena (in vitro-in vivo extrapolation (IVIVE) methods) to be incorporated, to obtain a more accurate result. Apart from the experiments and species conversion (for example, translation between the cell culture, animal and human), the challenge also includes the unique properties of nanomaterials that might cause them to behave differently compared to the same materials in a bulk form. This overview presents the IVIVE techniques that are developed to extrapolate pharmacokinetics data or doses. A brief example of the IVIVE methods for chemicals is provided, followed by a more detailed summary of available IVIVE methods applied to nanomaterials. The IVIVE techniques discussed include the comparison between in vitro and in vivo studies, methods to rene the dose metric or the in vitro models, allometric approach, mechanistic modeling, Multiple-Path Particle Dosimetry (MPPD), methods using organ burden data and also approaches that are currently being developed.
BACKGROUND:Changes in NR3C1 and IGF2/H19 methylation patterns have been associated with behavioural and psychiatric outcomes. Maternal mental state has been associated with offspring NR3C1 promotor and IGF2/H19 imprinting control region (ICR) methylation patterns. However, there is a lack of prospective studies with long-term follow-up. METHODS:52 mother-offspring pairs were studied from 12 to 22 weeks of pregnancy and offspring was followed-up until 28-29 years-of-age. During pregnancy, mothers filled in a Life Event Scale and a Daily Hassles Scale measuring perceived stress; i.e., appraisal or subjectively experienced severity of impact of important life events and of daily hassles in several life domains during pregnancy, respectively. Green space was quantified around the residence, using high-resolution (1 m2) map data. Saliva and blood samples were obtained from the adult offspring. Absolute DNA methylation levels were determined in blood and saliva on four NR3C1 amplicons, and one IGF2/H19 ICR amplicon using a bisulfite PCR and sequencing method. Linear mixed effect models were used to test the associations between perceived stress and green spaces during pregnancy, and adult offspring methylation patterns. RESULTS:We found associations between maternal perceived stress during pregnancy and methylation patterns on two out of the four NR3C1 amplicons, measured in blood, from offspring in adulthood, but not with IGF2/H19 methylation. For an interquartile-range (IQR) increase in maternal perceived life event or daily hassles stress scores, absolute methylation levels on several NR3C1 CpG sites were significantly changed (-1.62 % to +5.89 %, p<0.05). Maternal perceived stress scores were not associated with IGF2/H19 methylation, neither in blood nor in saliva. Maternal exposure to green spaces surrounding the residence during the pregnancy was associated with IGF2/H19 ICR methylation (-0.80 % to -1.04 %, p<0.05) in saliva, but not with NR3C1 promotor methylation. CONCLUSION:We observed significant long-term effects of maternal perceived stress during pregnancy on the methylation patterns of the NR3C1 promotor in offspring well into adulthood. This may imply that maternal psychological distress during pregnancy may influence the regulation of the HPA-axis well into adulthood. Additionally, maternal proximity to green spaces was associated with IGF2/H19 ICR methylation patterns, which is a novel finding.
The application of nanomaterials in various fields has both benefits and risks. While the unique physico-chemical properties of nanomaterials are valuable for diverse applications, they may also alter how nanomaterials are taken up in the body, thus affecting their potential toxicity to humans. To ensure their safety, the potential risk of nanomaterials needs to be assessed. The risk of a given nanomaterial is typically evaluated through in vitro and in vivo studies. The H2020 NanoInformaTIX project, that is conducted under the European Union’s Horizon 2020 research and innovation programme is focusing on the study of nanomaterials. One of its objectives includes the identification of nanomaterial toxicity, of which the first step consists of the detection of a nanomaterial dose-response relationship. In this paper, we focus on this step, which can be done for a given nanomaterial toxicity endpoint of interest, using statistical tests for a monotone dose-response trend, i.e., monotone relationship between the increasing concentration of the nanomaterial and the toxicity endpoint. The significance of the monotone trend is tested using a likelihood ratio test and t-tests. An R package and a Shiny App were developed to perform the analysis. As an illustration, data extracted from the H2020 NanoInformaTIX platform is presented as a case study. This data consists of genetic toxicity studies with DNA strand breaks as the toxicity endpoint of interest. Significant trends were found for 12 out of 26 nanomaterials after adjusting for multiplicity by applying the False discovery rate (FDR) method.
Identification and isolation of COVID-19 infected persons plays a significant role in the control of COVID-19 pandemic. A country's COVID-19 positive testing rate is useful in understanding and monitoring the disease transmission and spread for the planning of intervention policy. Using publicly available data collected between March 5th, 2020 and May 31st, 2021, we proposed to estimate both the positive testing rate and its daily rate of change in South Africa with a flexible semi-parametric smoothing model for discrete data. There was a gradual increase in the positive testing rate up to a first peak rate in July, 2020, then a decrease before another peak around mid-December 2020 to mid-January 2021. The proposed semi-parametric smoothing model provides a data driven estimates for both the positive testing rate and its change. We provide an online R dashboard that can be used to estimate the positive rate in any country of interest based on publicly available data. We believe this is a useful tool for both researchers and policymakers for planning intervention and understanding the COVID-19 spread.
Parameter estimation is often considered as a post model selection problem, i.e., the parameters of interest are often estimated based on “the best” model. However, this approach does not take into account that “the best” model was selected from a set of possible models. Ignoring this uncertainty may lead to bias in estimation. In this paper, we present a Bayesian variable selection (BVS) approach for model averaging which would address the model uncertainty. Although averaging would be preferred approach, BVS can be used as well for model selection if the interest is to select one among the set of candidate models. The performance of Bayesian variable selection is compared with the information criterion based model averaging on real longitudinal data and through simulations study.
Abstract Funding Acknowledgements Type of funding sources: None. Background Telemonitoring is an intervention that has shown to improve care of heart failure (HF) patients and thus plays a major role in preventing frequent hospital visits. Purpose This study examined the impact of HF telemonitoring on unplanned hospitalization. Methods Patients admitted to the hospital for HF and who agreed to use the non-invasive telemonitoring system were candidates in the study. It measured body weight, blood pressure and heart rate each day, and an alert was made if they exceeded a certain limit. The study contained data from 97 patients. The primary endpoint was a composite of unplanned cardiovascular hospitalization and all-cause mortality. Patients were classified into three groups: hospitalized (for any of the endpoints) during the telemonitoring period (Group 1 ["monitoring failure"]), hospitalized after the telemonitoring period (Group 2), and not hospitalized (Group 3). Results Median telemonitoring period was 222 [141, 518] days. The number of patients in each group was 18, 18, and 61. Patients in Group 1 were older (75 vs. 63 years, p < 0.001) and had poorer estimated glomerular filtration rate (31 vs. 62 mL/min/1.73m2, p < 0.001) than those in Group 3. Length of hospitalization periods at the endpoint tended to be longer in Group 1 than Group 2 (8.5 [6, 14] vs. 5 [4, 10] days, p = 0.084). Conclusion If heart failure telemonitoring fails to prevent unplanned hospitalization, the length of that hospitalization period may be longer than if the unplanned hospitalization could have been prevented.
Several major global challenges, including climate change and water scarcity, warrant a scientific approach to generating solutions. Developing high quality and robust capacity in (bio)statistics is key to ensuring sound scientific solutions to these challenges, so collaboration between academic and research institutes should be high on university agendas. To strengthen capacity in the developing world, South–North partnerships should be a priority. The ideas and examples of statistical capacity-building presented in this article are the result of several monthly online discussions between a mixedgroup of authors having international experience and formal links with Hasselt University in Belgium. The discussion focuses on statistical capacity-building through education (teaching), research, and societal impact. We have adopted an example-based approach, and in view of the background of the authors, the examples refer mainly to biostatistical capacity-building. Although many universities worldwide have already initiated university collaborations for development, we hope and believe that our ideas and concrete examples can serve as inspiration to further strengthen South–North partnerships on statistical capacity-building.
AIMS:There has been an increased interest in studying the association between microbial communities and different diseases and in discovering microbiome biomarkers. This association is pivotal to discover such biomarkers. In this paper, we present a unified modelling approach that can be used to detect and develop microbiome biomarkers for different clinical responses of interest at different levels of the microbiome ecosystem.METHODS AND RESULTS:We extended the methodology rooted in the information theory and joint modelling approaches for the evaluation of surrogate endpoints in randomized clinical trials to the high-dimensional microbiome setting. The unified modelling approach introduced in this paper allows for detecting biomarkers associated with a clinical response of interest, adjusting for the intervention applied to the subjects. For some microbiome features, the association is driven by the treatment, while for others, the association reflects the correlation between the microbiome biomarker and the clinical response of interest.CONCLUSIONS:The results have demonstrated that biomarkers can be identified at different levels of the microbiome phylogenetic tree using various measures as biomarkers.
Abstract Background In resource-limited settings, changes in CD4 counts constitute an important component in patient monitoring and evaluation of treatment response as these patients do not have access to routine viral load testing. In this study, we quantified trends on CD4 counts in patients on highly active antiretroviral therapy (HAART) in a comprehensive health care clinic in Kenya between 2011 and 2017. We evaluated the rate of change in CD4 cell count in response to antiretroviral treatment. We further assessed factors that influenced time to treatment change focusing on baseline characteristics of the patients and different initial drug regimens used. This was a retrospective study involving 432 naïve HIV patients that had at least two CD4 count measurements for the period. The relationship between CD4 cell count and time was modeled using a semi parametric mixed effects model while the Cox proportional hazards model was used to assess factors associated with the first regimen change. Results Majority of the patients were females and the average CD4 count at start of treatment was 362.1 $$cell/mm^3$$ c e l l / m m 3 . The CD4 count measurements increased nonlinearly over time and these trends were similar regardless of the treatment regimen administered to the patients. The change of logarithm CD4 cell count rises fast for in the first 450 days of antiretroviral initiation. The average time to first regimen change was 2142 days. Tenoforvir (TDF) based regimens had a lower drug substitution(aHR 0.2682, 95% CI:0.08263- 0.8706) compared to Zidovudine(AZT). Conclusion The backbone used was found to be associated with regimen changes among the patients with fewer switches being observed, with the use of TDF when compared to AZT. There was however no significant difference between TDF and AZT in terms of the rate of change in logarithm CD4 count over time.
Bayesian inference for generalized linear mixed models (GLMM) is appealing, but its widespread use has been hampered by the lack of a fast implementation tool and the difficulty in specifying prior distributions. In this paper, we conduct an extensive simulation study to evaluate the performance of INLA for estimation of the hierarchical Poisson regression models with overdispersion in comparison with JAGS and Stan while assuming a variety of prior specifications for variance components. Further, we analysed the influence of different factors such as small number of observations per cluster, different values of the cluster variance and estimation from a misspecified model. A simulation study has shown that the approximation strategy employed by INLA is accurate in general and that all software leads to similar results for most of the cases considered. Estimation of the variance components, however, is difficult when their true value is small for all estimation methods and prior specifications. The estimates obtained for all software tend to be biased downward or upward depending on the assumed priors.
INTRODUCTION:In contrast to most tuberculosis (TB) high burden countries, Ethiopia has for a long time reported a very high percentage of extra pulmonary TB (EPTB), which is also reflected in population based estimations reported by the World Health Organization (WHO). Particularly a steadily higher proportion of cervical tuberculous lymphadenitis (TBLN) has been described. Here we identify clinical and demographic factors associated with anatomic site of the TB disease. METHOD:A health facility based comparative study was conducted among TBLN and PTB patients who visited selected health facilities in Ethiopia during 2016 and 2017. Associated risk factors were identified through a multivariate logistic regression model using R-studio. RESULT:A total of 1,890 study participants, 427 TBLN and 1,463 PTB patients, were included. The mean age of TBLN patients (29 years ± 14.4 SD) was lower than that of PTB cases (36 years ± 15.0 SD). There were slightly more women diagnosed with TBLN (51.1%) while nearly 6 out of 10 male patients were diagnosed with PTB (58.9%). Most significantly, younger age groups (<15 Years) were more likely to develop cervical TBLN than older people (>56 years), with an AOR of 9.76 (95% CI: 4.87, 19.56). The odds of cervical TBLN among women [1.69 (1.30, 2.20)] was higher than that for men. In addition, adjusted estimates suggested that, compared with PTB, renal diseases [3.41 (1.29, 9.02)] and the presence of other concomitant chronic illness [1.61 (1.23, 2.09)] had a significant association with TBLN. CONCLUSION:Generally, the risk of developing a particular form of TB disease is usually associated with demographic and medical history of an infected individual. Hence, the current symptom based screening, which primarily rely on chronic cough in many countries, may lead to missing significant portions of TBLN cases.
Background In resource-limited settings, changes in CD4 counts constitute an important component in patient monitoring and evaluation of treatment response as these patients do not have access to routine viral load testing. In this study, we quantified trends on CD4 counts in patients on highly active antiretroviral therapy (HAART) in a comprehensive health care clinic in Kenya between 2011 and 2017.We evaluated the rate of change in CD4 cell count in response to antiretroviral treatment. We further assessed factors that influenced time to treatment change focusing on baseline characteristics of the patients and different initial drug regimens used. The study involved 529 naïve HIV patients that had at least two CD4 count measurements for the period. The relationship between CD4 cell count and time was modeled using a semi parametric mixed effects model while the Cox proportional hazards model was used to assess factors associated with the first regimen change. Results The results demonstrated that CD4 counts increased over time and these trends were similar regardless of the treatment regimen used. Males were less likely to have drug regimens switch (adjusted hazard ratio (aHR) 0.5101, 95% CI: 0.1906 -1.3647) compared to females.Tenoforvir (TDF) based regimens had a lower drug substitution (aHR 0.2796, 95% CI: 0.0961-0.8629) compared to Zinovudine (AZT). Conclusion The backbone used was found to be associated with regimen changes among the patients with fewer switches being observed, with the use of TDF when compared to AZT. There was however no significant difference between TDF and AZT in terms of the change in CD4 count over time.
Our aim in this study is to develop predictive microbiome biomarkers for intestinal IgA levels. In this article, a operational taxonomic units(OTU)-specific (family-specific) and time-specific joint model is presented as a tool to model the association between OTU (or family) and biological response (measured by IgA level) taking into account the treatment group (Control or PAT) of the subjects. The model allows detecting OTUs (families) that are associated with the IgA; for some OTUs (families), the association is driven by the treatment while for others the association reflects the correlation between the OTUs (families) and IgA.The results of the analysis reveal that: (1) the observed diversity of S24-7 family can be used as a biomarker to classify samples according to treatment group for days 6 and 12; (2) the treatment effect induces the corrlelation between the S24-7 diversity and the IgA level at day 20; (3) The OTUs that are identified to be significantly differentially abundant (FDR level of 0.05) between the two treatment groups for days 12 and 20 are all part of the S24-7 family, although most of the differentially abundant ones at day 1 are from the Lactobacillaceae family; (4) only the Lachnospiraceae family diversity at day 6, and 20 can be used as predictive biomarker for the IgA level at day 20; (5) New.ReferenceOTU513, correlated with the IgA level at day 20, since day 12, belongs to the Lachnospiraceae family and all other OTUs among the top 10 significantly associated OTUs at day 20 are from the S24-7 family; (6) the observed alpha diversity at day 6 is significantly differentially abundant and can be used as predictive biomarker for IgA level at day 20.
Background Previous work has shown differential predominance of certain Mycobacterium tuberculosis (M . tb) lineages and sub-lineages among different human populations in diverse geographic regions of Ethiopia. Nevertheless, how strain diversity is evolving under the ongoing rapid socio-economic and environmental changes is poorly understood. The present study investigated factors associated with M . tb lineage predominance and rate of strain clustering within urban and peri-urban settings in Ethiopia. Methods Pulmonary Tuberculosis (PTB) and Cervical tuberculous lymphadenitis (TBLN) patients who visited selected health facilities were recruited in the years of 2016 and 2017. A total of 258 M . tb isolates identified from 163 sputa and 95 fine-needle aspirates (FNA) were characterized by spoligotyping and compared with international M . tb spoligotyping patterns registered at the SITVIT2 databases. The molecular data were linked with clinical and demographic data of the patients for further statistical analysis. Results From a total of 258 M . tb isolates, 84 distinct spoligotype patterns that included 58 known Shared International Type (SIT) patterns and 26 new or orphan patterns were identified. The majority of strains belonged to two major M . tb lineages, L3 (35.7%) and L4 (61.6%). The observed high percentage of isolates with shared patterns (n = 200/258) suggested a substantial rate of overall clustering (77.5%). After adjusting for the effect of geographical variations, clustering rate was significantly lower among individuals co-infected with HIV and other concomitant chronic disease. Compared to L4, the adjusted odds ratio and 95% confidence interval (AOR; 95% CI) indicated that infections with L3 M . tb strains were more likely to be associated with TBLN [3.47 (1.45, 8.29)] and TB-HIV co-infection [2.84 (1.61, 5.55)]. Conclusion Despite the observed difference in strain diversity and geographical distribution of M . tb lineages, compared to earlier studies in Ethiopia, the overall rate of strain clustering suggests higher transmission and warrant more detailed investigations into the molecular epidemiology of TB and related factors.