We consider the problem of relating two high-dimensional datasets with the goal of identifying subsets of outcomes with the same regression function over a subset of covariates, allowing for nonlinear relationships. This is accomplished by specifying a mixture of Gaussian process regression models and performing variable selection using a stochastic partitioning method. The proposed method provides a flexible approach that simultaneously uncovers cluster structures and identifies linear and nonlinear relationships across two large datasets. We evaluate its performance on simulated data in terms of identification of cluster structures, variable selection, and prediction.
A graph structure is commonly used to characterize the dependence between variables, which may be induced by time, space, biological networks or other factors. Incorporating this dependence structure into the variable selection procedure can improve the identification of relevant variables, especially those with subtle effects. The Bayesian approach provides a natural framework to integrate this information through the prior distributions. In this work, we propose combining two priors that have been well studied separately-the Gaussian Markov random field prior and the horseshoe prior-to perform selection on graph-structured variables. Local shrinkage parameters that capture the dependence between connected covariates are specified to encourage similar amount of shrinkage for their regression coefficients, while a standard horseshoe prior is used for non-connected variables. After evaluating the performance of the method on different simulated scenarios, we present three applications: one in quantitative trait loci mapping with block sequential structure, one in near-infrared spectroscopy with sequential non-disjoint dependence and another in gene expression study with a general dependence structure.
SummaryThe understanding of tree growth processes is crucial for promoting sustainable forest management strategies. This is a challenging task in highly biodiverse ecosystems where many tree species are observed on very few individuals and the small sample sizes hinder a good fit of species‐specific models. We propose the use of finite mixture of random coefficient regression models with multilevel nested random effects to infer guild specific fixed and random effects while evaluating the relative importance of the nested sources of variability on goodness‐of‐fit. This approach extends finite mixture of linear mixed model used for longitudinal or single group structured data contexts. A dedicated expectation–maximisation algorithm is introduced for parameter estimation. Simulations are performed for the evaluation of the misspecification of nested‐grouping structures. This work has been motivated by data collected biennially in Central African rainforests from 1986 to 2010. We show the accuracy of the proposed approach in successfully reproducing individual growth processes and classifying tree species into well‐differentiated clusters with clear ecological interpretations. Moreover, results confirm that interindividual variability appears as the most important factor to explain tropical tree species growth process variability from Central Africa forests.
Oncogenesis, a complex and multifaceted process, is profoundly modulated by miRNA’s regulatory role in gene expression. Over the years, a substantial body of knowledge concerning miRNA and mRNA has been accumulated, drawing from both rigorous biological experiments and intricate statistical analyses. In the realm of statistical modeling, the integration of such information as "prior knowledge" often amplifies the model’s ability to pinpoint molecular targets of significance. This study seeks to leverage prior knowledge of miRNA-mRNA regulatory interactions to map the dynamic landscape of interactions in the specific context of hepatocellular carcinoma (HCC).To address this, we introduce an evolved iteration of a Bayesian two-step integrative method previously established in the literature. This augmented approach includes improved computing efficiency when dealing with high dimensional data and a novel mechanistic submodel, which operates autonomously, devoid of prior knowledge. Employing this method, we identified two discrete gene lists: one informed by prior knowledge and the other independently inferred. This bifurcated strategy provides a comprehensive perspective on gene interactions.Our methodological advancement allows for a nuanced analysis of gene networks, distinguishing between direct and indirect gene relationships and considering miRNA influences with two available sub-mechanistic submodels. We introduce an approach to validate our findings using a biological interaction network, emphasizing the quality and relevance of identified gene-gene relationships. Metrics like the Matthews Correlation Coefficient (MCC) and the true discovery rate (TDR) further attest to the robustness of our findings.In summation, aside from improving the existing sub-mechanistic model that requires prior knowledge, this paper presents an innovative prior knowledge-free sub-mechanistic model as an alternative. It champions the use of biological networks for validation, underscoring the significance of methodological advancements in genomics research.
The link between cofactor binding and protein activity is well-established. However, how cofactor interactions modulate folding of large proteins remains unknown. We use optical tweezers, clustering and global fitting to dissect the folding mechanism of Drosophila cryptochrome (dCRY), a 542-residue protein that binds FAD, one of the most chemically and structurally complex cofactors in nature. We show that the first dCRY parts to fold are independent of FAD, but later steps are FAD-driven as the remaining polypeptide folds around the cofactor. FAD binds to largely unfolded intermediates, yet with association kinetics above the diffusion-limit. Interestingly, not all FAD moieties are required for folding: whereas the isoalloxazine ring linked to ribitol and one phosphate is sufficient to drive complete folding, the adenosine ring with phosphates only leads to partial folding. Lastly, we propose a dCRY folding model where regions that undergo conformational transitions during signal transduction are the last to fold.
Abstract The functional link between cofactor binding and protein activity is well established, but how cofactor interactions with the polypeptide reshape the folding energy landscape of large and multidomain proteins is unknown. Here, we use optical tweezers in combination with a novel analytical framework that integrates clustering, bootstrapping and global fitting of kinetic and thermodynamic data to dissect the folding mechanism of the light-sensing Drosophila cryptochrome (dCRY), a 542-residue protein that binds FAD, one of the most common, complex cofactors. We show that FAD binds to multiple dCRY folding intermediates, some of which contain large amounts of unfolded polypeptide. Yet, binding occurs with association kinetics above the diffusion-limit and at sub-nanomolar affinity. Surprisingly, the first parts of dCRY to fold are independent of FAD, but later steps are FAD-driven as the remaining protein folds around the cofactor. Thus, dCRY coordinates cofactor-dependent and independent folding mechanisms to attain its native state. Furthermore, we find that not all the FAD chemical moieties are strictly required for folding: whereas the isoalloxazine ring linked to ribitol and one phosphate group (i.e., FMN) are sufficient to drive complete dCRY folding, the adenosine ring plus the phosphate groups (i.e., AMP and ADP) only allow partially folded structures. Lastly, by combining the results from optical tweezers experiments with structural data, we propose a model for the dCRY folding pathway wherein regions known to undergo conformational transitions during signal transduction are the last to fold. Altogether, our single-molecule experiments and data analysis illustrate the power and broad applicability of optical tweezers to dissect complex mechanisms that couple the folding of large proteins to cofactor binding.
Recent studies have confirmed the role of miRNA regulation of gene expression in oncogenesis for various cancers. In parallel, prior knowledge about relationships between miRNA and mRNA have been accumulated from biological experiments or statistical analyses. Improved identification of disease-associated miRNA-mRNA pairs may be achieved by incorporating prior knowledge into integrative genomic analyses. In this study we focus on 39 patients with hepatocellular carcinoma (HCC) and 25 patients with liver cirrhosis and use a flexible Bayesian two-step integrative method. We found 66 significant miRNA-mRNA pairs, several of which contain molecules that have previously been identified as potential biomarkers. These results demonstrate the utility of the proposed approach in providing a better understanding of relationships between different biological levels, thereby giving insights into the biological mechanisms underlying the diseases, while providing a better selection of biomarkers that may serve as diagnostic, prognostic, or therapeutic biomarker candidates.
Dissecting protein folding pathways is challenging, especially for large or multidomain proteins. This challenge is further exacerbated when considering that many proteins harbor one or more cofactors in their structure, which can alter their fold and thermodynamic properties. Here, we use optical tweezers to mechanically unfold and refold drosophila cryptochrome (dCRY), a large, multidomain protein that harbors a FAD cofactor in its structure. By applying force and simultaneously varying the FAD concentration and the refolding and rebinding dwell time, we dissect the mechanisms that coupled dCRY folding and FAD binding. We show that dCRY cannot fold into its native state in the absence of FAD. In contrast, FAD enables dCRY to fully fold in a stepwise fashion via five intermediates. Interestingly, FAD binds dCRY very fast (near the diffusion-limit), with sub-nanomolar affinity, and occurs at two intermediate states, possibly following two different binding mechanisms. By using a variety of cofactors that contain part of the chemical moieties of FAD, we find that the isoalloxazine ring linked to ribitol and one phosphate group (i.e., FMN) is the main driver of dCRY folding. The absence of the phosphate in FMN (i.e., riboflavin) significantly reduces the probability of dCRY to fold into its native state, underscoring the role of phosphate groups in folding. However, the phosphate groups linked to the adenosine ring of FAD (i.e., AMP or ADP) do not allow dCRY to attain its native state, which highlights the non-additive contributions of different chemical moieties in driving folding. Finally, by combining the results from optical tweezers experiments with structural data, we propose a model for the folding pathway coupled to cofactor binding for this protein.
Bayesian statistics is an approach to data analysis based on Bayes' theorem, where available knowledge about parameters in a statistical model is updated with the information in observed data. The background knowledge is expressed as a prior distribution and combined with observational data in the form of a likelihood function to determine the posterior distribution. The posterior can also be used for making predictions about future events. This Primer describes the stages involved in Bayesian analysis, from specifying the prior and data models to deriving inference, model checking and refinement. We discuss the importance of prior and posterior predictive checking, selecting a proper technique for sampling from a posterior distribution, variational inference and variable selection. Examples of successful applications of Bayesian analysis across various research fields are provided, including in social sciences, ecology, genetics, medicine and more. We propose strategies for reproducibility and reporting standards, outlining an updated WAMBS (when to Worry and how to Avoid the Misuse of Bayesian Statistics) checklist. Finally, we outline the impact of Bayesian analysis on artificial intelligence, a major goal in the next decade. This Primer on Bayesian statistics summarizes the most important aspects of determining prior distributions, likelihood functions and posterior distributions, in addition to discussing different applications of the method across disciplines.
Background: Adjuvant endocrine therapy (AET) improves outcomes in women with hormone receptor–positive (HR+) breast cancer. Suboptimal AET adherence is common, but data are lacking about symptoms and adherence in racial/ethnic minorities. We evaluated adherence by race and the relationship between symptoms and adherence. Methods: The Women's Hormonal Initiation and Persistence study included women diagnosed with nonrecurrent HR+ breast cancer who initiated AET. AET adherence was captured using validated items. Data regarding patient (e.g., race), medication-related (e.g., symptoms), cancer care delivery (e.g., communication), and clinicopathologic factors (e.g., chemotherapy) were collected via surveys and medical charts. Multivariable logistic regression models were employed to calculate odds ratios and 95% confidence intervals (CIs) associated with adherence. Results: Of the 570 participants, 92% were privately insured and nearly one of three were Black. Thirty-six percent reported nonadherent behaviors. In multivariable analysis, women less likely to report adherent behaviors were Black (vs. White; OR, 0.43; 95% CI, 0.27–0.67; P < 0.001) and with greater symptom burden (OR, 0.98; 95% CI, 0.96–1.00; P < 0.05). Participants more likely to be adherent were overweight (vs. normal weight) (OR, 1.58; 95% CI, 1.04–2.43; P < 0.05), sat ≤ 6 hours a day (vs. ≥6 hours; OR, 1.83; 95% CI, 1.25–2.70; P < 0.01), and were taking aromatase inhibitors (vs. tamoxifen; OR, 1.91; 95% CI, 1.28–2.87; P < 0.01). Conclusions: Racial differences in AET adherence were observed. Longitudinal assessments of symptom burden are needed to better understand this dynamic process and factors that may explain differences in survivor subgroups. Impact: Future interventions should prioritize Black survivors and women with greater symptom burden.
A Correction to this paper has been published: https://doi.org/10.1038/s43586-021-00017-2.
Pathologic alterations in epigenetic regulation have long been considered a hallmark of many cancers, including hepatocellular carcinoma (HCC). In a healthy individual, the relationship between DNA methylation and microRNA (miRNA) expression maintains a fine balance; however, disruptions in this harmony can aid in the genesis of cancer or the propagation of existing cancers. The balance between DNA methylation and microRNA expression and its potential disturbance in HCC can vary by race. There is emerging evidence linking epigenetic events including DNA methylation and miRNA expression to cancer disparities. In this paper, we evaluate the epigenetic mechanisms of racial heterogenity in HCC through an integrated analysis of DNA methylation, miRNA, and combined regulation of gene expression. Specifically, we generated DNA methylation, mRNA-seq, and miRNA-seq data through the analysis of tumor and adjacent non-tumor liver tissues from African Americans (AA) and European Americans (EA) with HCC. Using mixed ANOVA, we identified cytosine-phosphate-guanine (CpG) sites, mRNAs, and miRNAs that are significantly altered in HCC vs. adjacent non-tumor tissue in a race-specific manner. We observed that the methylome was drastically changed in EA with a significantly larger number of differentially methylated and differentially expressed genes than in AA. On the other hand, the miRNA expression was altered to a larger extent in AA than in EA. Pathway analysis functionally linked epigenetic regulation in EA to processes involved in immune cell maturation, inflammation, and vascular remodeling. In contrast, cellular proliferation, metabolism, and growth pathways are found to predominate in AA as a result of this epigenetic analysis. Furthermore, through integrative analysis, we identified significantly differentially expressed genes in HCC with disparate epigenetic regulation, associated with changes in miRNA expression for AA and DNA methylation for EA.
A Correction to this paper has been published: https://doi.org/10.1038/s43586-021-00017-2.
Background: The established role miRNA-mRNA regulation of gene expression has in oncogenesis highlights the importance of integrating miRNA with downstream mRNA targets. These findings call for investigations aimed at identifying disease-associated miRNA-mRNA pairs. Hierarchical integrative models (HIM) offer the opportunity to uncover the relationships between disease and the levels of different molecules measured in multiple omic studies. Methods: The HIM model we formulated for analysis of mRNA-seq and miRNA-seq data can be specified with two levels: (1) a mechanistic submodel relating mRNAs to miRNAs, and (2) a clinical submodel relating disease status to mRNA and miRNA, while accounting for the mechanistic relationships in the first level. Results: mRNA-seq and miRNA-seq data were acquired by analysis of tumor and normal liver tissues from 30 patients with hepatocellular carcinoma (HCC). We analyzed the data using HIM and identified 157 significant miRNA-mRNA pairs in HCC. The majority of these molecules have already been independently identified as being either diagnostic, prognostic, or therapeutic biomarker candidates for HCC. These pairs appear to be involved in processes contributing to the pathogenesis of HCC involving inflammation, regulation of cell cycle, apoptosis, and metabolism. For further evaluation of our method, we analyzed miRNA-seq and mRNA-seq data from TCGA network. While some of the miRNA-mRNA pairs we identified by analyzing both our and TCGA data are previously reported in the literature and overlap in regulation and function, new pairs have been identified that may contribute to the discovery of novel targets. Conclusion: The results strongly support the hypothesis that miRNAs are important regulators of mRNAs in HCC. Furthermore, these results emphasize the biological relevance of studying miRNA-mRNA pairs.
Abstract Purpose: Adjuvant endocrine therapy (AET) improves survival in women with hormone receptor-positive (HR+) breast cancer (BC). Yet medication adherence is suboptimal. The aim of this study was to assess adherence to AET among insured women using innovative statistical approaches. Methods: Black and White women diagnosed with HR+ BC were identified from two health maintenance organizations. Automated pharmacy records captured oral AET prescriptions and refill dates. Logistic regression identified predictors of adherence defined in terms of proportion of days covered (PDC) (>=80%) and medication gap of ≤10 days. A zero-inflated negative binominal (ZINB) regression model identified variables associated with the total number of days of medication gaps. Results: A total of 1,925 women met inclusion criteria. Eighty percent of women were adherent per the PDC measure; 44% had a medication gap of ≤10 days; and 24% of women had zero days without any medication gaps. Race and age were significant predictors of adherence in all multivariable models. Black women were less likely to have PDC >=80% than Whites (OR=0.72; 95%CI: 0.57-0.90; p<0.01), and they were less likely to have a medication gap of ≤10 days (OR=0.65; 95%CI: 0.54-0.79; p<0.001). Women 25-49 years old were less likely to have PDC >=80% than women 65-93 years old (OR=0.65; 95%CI: 0.48-0.87; p<0.001), and they also were less likely to have a medication gap of ≤10 days (OR=0.73; 95%CI: 0.57-0.93; p<0.01). In the zero-inflated negative binominal model, Black women were less likely to having no medication gaps compared to Whites (OR=0.46; 95%CI: 0.54-0.79; p<0.001), and women 25-49 years old were less likely to have no medication gaps compared to women 65-93 years old (OR=0.61; 95%CI: 0.42-0.88; p<0.01). Conclusions: Disparities in adherence to AET persist among insured women, particularly in Black and young women, highlighting a need for interventions among this population. Novel statistical approaches to study adherence, such as the ZINB approach, appear to constitute a useful alternative to the dichotomous PDC variable to tailor analysis to adherence patterns. Citation Format: Dennis Tolsma, Mahlet G. Tadesse, Arnethea Sutton, Lee Cromwell, Georges Adunlin, Teresa M. Salgado, Jun He, Martha Trout, Brandi E. Robinson, Megan C. Edmonds, Hayden B. Bosworth, Vanessa B. Sheppard. Adherence to adjuvant endocrine therapy: Do racial disparities persist among the insured? [abstract]. In: Proceedings of the Eleventh AACR Conference on the Science of Cancer Health Disparities in Racial/Ethnic Minorities and the Medically Underserved; 2018 Nov 2-5; New Orleans, LA. Philadelphia (PA): AACR; Cancer Epidemiol Biomarkers Prev 2020;29(6 Suppl):Abstract nr A076.
BACKGROUND:The onset of silent diseases such as type 2 diabetes is often registered through self-report in large prospective cohorts. Self-reported outcomes are cost-effective; however, they are subject to error. Diagnosis of silent events may also occur through the use of imperfect laboratory-based diagnostic tests. In this paper, we describe an approach for variable selection in high dimensional datasets for settings in which the outcome is observed with error.METHODS:We adapt the spike and slab Bayesian Variable Selection approach in the context of error-prone, self-reported outcomes. The performance of the proposed approach is studied through simulation studies. An illustrative application is included using data from the Women's Health Initiative SNP Health Association Resource, which includes extensive genotypic (>900,000 SNPs) and phenotypic data on 9,873 African American and Hispanic American women.RESULTS:Simulation studies show improved sensitivity of our proposed method when compared to a naive approach that ignores error in the self-reported outcomes. Application of the proposed method resulted in discovery of several single nucleotide polymorphisms (SNPs) that are associated with risk of type 2 diabetes in a dataset of 9,873 African American and Hispanic participants in the Women's Health Initiative. There was little overlap among the top ranking SNPs associated with type 2 diabetes risk between the racial groups, adding support to previous observations in the literature of disease associated genetic loci that are often not generalizable across race/ethnicity populations. The adapted Bayesian variable selection algorithm is implemented in R. The source code for the simulations are available in the Supplement.CONCLUSIONS:Variable selection accuracy is reduced when the outcome is ascertained by error-prone self-reports. For this setting, our proposed algorithm has improved variable selection performance when compared to approaches that neglect to account for the error-prone nature of self-reports.
Maternal genetic variations, including variations in mitochondrial biogenesis (MB) and oxidative phosphorylation (OP), have been associated with placental abruption (PA). However, the role of maternal-fetal genetic interactions (MFGI) and parent-of-origin (imprinting) effects in PA remain unknown. We investigated MFGI in MB-OP, and imprinting effects in relation to risk of PA. Among Peruvian mother-infant pairs (503 PA cases and 1,052 controls), independent single nucleotide polymorphisms (SNPs), with linkage-disequilibrium coefficient <0.80, were selected to characterize genetic variations in MB-OP (78 SNPs in 24 genes) and imprinted genes (2713 SNPs in 73 genes). For each MB-OP SNP, four multinomial models corresponding to fetal allele effect, maternal allele effect, maternal and fetal allele additive effect, and maternal-fetal allele interaction effect were fit under Hardy-Weinberg equilibrium, random mating, and rare disease assumptions. The Bayesian information criterion (BIC) was used for model selection. For each SNP in imprinted genes, imprinting effect was tested using a likelihood ratio test. Bonferroni corrections were used to determine statistical significance (p-value<6.4e-4 for MFGI and p-value<1.8e-5 for imprinting). Abruption cases were more likely to experience preeclampsia, have shorter gestational age, and deliver infants with lower birthweight compared with controls. Models with MFGI effects provided improved fit than models with only maternal and fetal genotype main effects for SNP rs12530904 (log-likelihood ratio=18.2; p-value=1.2e-04) in CAMK2B , and, SNP rs73136795 (log-likelihood ratio=21.7; p-value=1.9e-04) in PPARG , both MB genes. We identified 311 SNPs in 35 maternally-imprinted genes (including KCNQ1, NPM , and, ATP10A) associated with abruption. Top hits included rs8036892 (p-value=2.3e-15) in ATP10A , rs80203467 (p-value=6.7e-15) and rs12589854 (p-value=1.4e-14) in MEG8 , and rs138281088 in SLC22A2 (p-value=1.7e-13). We identified novel PA-related maternal-fetal MB gene interactions and imprinting effects that highlight the role of the fetus in PA risk development. Findings can inform mechanistic investigations to understand the pathogenesis of PA. Author summary Placental Abruption (PA) is a complex multifactorial and heritable disease characterized by premature separation of the placenta from the wall of the uterus. PA is a consequence of complex interplay of maternal and fetal genetics, epigenetics, and metabolic factors. Previous studies have identified common maternal single nucleotide polymorphisms (SNPs) in several mitochondrial biogenesis (MB) and oxidative phosphorylation (OP) genes that are associated with PA risk, although findings were inconsistent. Using the largest assembled mother-infant dyad of PA cases and controls, that includes participants from a previous report, we identified novel PA-related maternal-fetal MB gene interactions and imprinting effects that highlight the role of the fetus in PA risk development. Our findings have the potential for enhancing our understanding of genetic variations in maternal and fetal genome that contribute to PA.
In addition to socioeconomic influences, biological factors are believed to play a role in health disparities. In this paper, we investigate miRNA, mRNA, and DNA methylation patterns that contribute to disparities in hepatocellular carcinoma (HCC). This is accomplished by integration of mRNA-Seq, miRNA-Seq, and DNA methylation data we acquired by analysis of liver tissues from 30 HCC patients consisting of European Americans (EAs), African Americans (AAs), and Asian Americans (Asians). Mixed-ANOVA models are applied to identify miRNAs, mRNAs, and DNA methylation sites that are significantly altered in tumor vs. adjacent normal tissues in a race-specific manner. Through integrated analysis, a refined list of differentially expressed mRNAs is obtained by selecting those that are targets of differentially expressed miRNAs and consist of promoter regions that are differentially methylated.
BACKGROUND:Adjuvant endocrine therapy (AET) is a critical therapy in that it improves survival in women with hormone receptor-positive (HR+) breast cancer (BC), but adherence to AET is suboptimal. The purpose of this study was to fill scientific gaps about predictors of adherence to AET among black and white women diagnosed with BC.OBJECTIVE:To assess AET adherence in black and white insured women using multiple measures, including one that uses an innovative statistical approach.METHODS:Black and white women newly diagnosed with HR+ BC were identified from 2 health maintenance organizations. Pharmacy records captured the type of oral AET prescriptions and all fill dates. Multivariable logistic regression was used to identify predictors of adherence defined in terms of proportion of days covered (PDC; ≥ 80%) and medication gap of ≤ 10 days. A zero-inflated negative binomial (ZINB) regression model was used to identify variables associated with the total number of days of medication gaps.RESULTS:1,925 women met inclusion criteria. 80% were PDC adherent (> 80%); 44% had a medication gap of ≤ 10 days; and 24% had no medication gap days. Race and age were significant in all multivariable models. Black women were less likely to be adherent based on PDC than white women (OR = 0.72, 95% CI = 0.57-0.90, P < 0.01), and they were less likely to have a medication gap of ≤ 10 days (OR = 0.65, 95% CI = 0.54-0.79, P < 0.001). Women aged 25-49 years were less likely to be PDC adherent than women aged 65-93 years (OR = 0.65, 95% CI = 0.48-0.87, P < 0.001). In the ZINB model, women were without their medication for an average of 37 days (SD = 50.5).CONCLUSIONS:Racial disparities in adherence to AET in the study highlight a need for interventions among insured women. Using various measures of adherence may help better understand this multidimensional concept. There might be benefits from using both more common dichotomous measures (e.g., PDC) and integrating novel statistical approaches to allow tailoring adherence to patterns within a specific sample.DISCLOSURES:This research was funded by the National Institutes of Health (R01CA154848). It was also supported in part by the NIH-NCI Cancer Center Support Grant P30 CA016059, the Laboratory of Telomere Health P30 CA51008, and the TSA Award No. UL1TR002649 from the National Center for Advancing Translational Sciences. The contents of this study are solely the responsibility of the authors and do not necessarily represent official views of the National Center for Advancing Translational Sciences or the National Institutes of Health. Bosworth reports grants from Sanofi, Otsuka, Johnson & Johnson, and Blue Cross/Blue Shield of NC and consulting fees from Sanofi and Otsuka. The other authors have nothing to disclose. The datasets generated during and/or analyzed during the current study are not publicly available due to privacy reasons but are available from the corresponding author on reasonable request. The author does not own these data. Data use was granted to the author as part of a data use agreement between specific agencies and organizations.