This work shows how Gaussian Graphical Models (GGMs) varying on a spatial network offer an effective means to accurately portray the conditional dependence relationships among water pollutants observed on a fluvial network. The motivating case study involves evaluating the quality of water bodies in the Piedmont region. In environmental sciences, representing variable relationships through estimated graphs proves both practical and advantageous. After estimating all graphs over a spatial network, it will be possible to identify spatially varying clusters of variables across the fluvial network
BACKGROUND:The association between allergic sensitization and asthma is well-documented, but its precise role in asthma remains uncertain. Component-resolved diagnostics allows detailed assessment of IgE-sensitization to multiple allergenic molecules (c-sIgE). We applied advanced network embedding techniques to investigate the dynamics of temporal development of multiple c-sIgE and identify networks associated with asthma. METHODS:In a population-based birth cohort, we measured c-sIgE to 112 proteins using multiplex array at 6 time points from infancy to adolescence. We built weighted co-occurrence networks between c-sIgEs to investigate connectivity structures at different ages. To identify critical periods where networks are similar/divergent, we applied graph embedding and dimensionality reduction techniques. We then compared network development structure between subjects with and without asthma at different ages and analyzed topological features to compare network structures. RESULTS:c-sIgE sensitization networks across ages revealed significant changes and a continuous evolution rather than abrupt shifts, with networks at ages 5 and 8 being very similar. Individuals with asthma consistently exhibited more complex and interconnected networks of c-sIgEs, which became more pronounced with age. Graph embedding showed that profiles of those with and without asthma were distinct and the separation persisted across ages. A specific set of c-sIgEs and their interactions were responsible for this distinction. Topological features of networks that distinguished between sensitized individuals with and without asthma were age-dependent. CONCLUSIONS:The differences in c-sIgE networks between subjects with and without asthma are consistently observed throughout childhood. Age needs to be considered when developing interpretation algorithms for asthma diagnosis/prediction.
BACKGROUND:Many studies used information on wheeze presence/absence to determine asthma-related phenotypes. We investigated whether clinically intuitive asthma subtypes can be identified by applying data-driven semi-supervised techniques to information on frequency and triggers of different respiratory symptoms. METHODS:Partitioning Around Medoids clustering was applied to data on multiple symptoms and their triggers in school-age children from three birth cohorts: MAAS (n = 947, age 8 years), SEATON (n = 763, age 10) and ASHFORD (n = 584, age 8). 'Guided' clustering, incorporating asthma diagnosis, was used to select the optimal number of clusters. RESULTS:Five-cluster solution was optimal. Based on their clinical characteristics, including frequency of asthma diagnosis, we interpreted one cluster as 'Healthy'. Two clusters were characterised by high asthma prevalence (95.89% and 78.13%). We assigned children with asthma in these two clusters as 'persistent, multiple-trigger, more severe' (PMTS) and 'persistent, triggered by infection, milder' (PIM). Children with asthma in the remaining two clusters were assigned as 'mild-remitting wheeze' (MRW) and 'post-bronchiolitis resolving asthma' (PBRA). PBRA was associated with RSV bronchiolitis in infancy. In most children with asthma in this cluster wheezing resolved by age 5-6, and predominant symptoms were shortness of breath and chest tightness. Children in PBRA had the highest hospitalisation rates and wheeze exacerbations in infancy. From age 8 years (cluster derivation) to early adulthood (18-20 years), lung function was significantly lower, and FeNO and airway hyperreactivity significantly higher in PMTS compared to all other clusters. CONCLUSIONS:Patterns of coexisting symptoms identified by semi-supervised data-driven methods may reflect pathophysiological mechanisms of distinct subtypes of childhood wheezing disorders.
Asthma is a heterogeneous condition often studied through wheeze alone, yet the interplay between lung function and reported symptoms remains underexplored. To capture this heterogeneity, we applied Bayesian Profile Regression to data from school-age children in two prospective birth cohorts, integrating airway hyperresponsiveness, lung function, bronchodilator reversibility, allergic sensitisation, reported symptoms, and physician diagnosis. In the Manchester Allergy and Asthma Study (discovery cohort), five reproducible clusters were identified: HA-LLF (high asthma-low lung function), HA-NLF (high asthma normal lung function), LA-RLF (low asthma-reduced lung function), LA-NLF (low asthma normal lung function), and MA-NLF (moderate asthma normal lung function). The HA-LLF and HA-NLF clusters had very high asthma prevalence (80–100
The emergence of big data and analytic approaches initiated research efforts to characterise different subtypes of allergic diseases, including tracking disease progression and identifying patterns that may offer insight into their development and progression. Triangulation from different data sources and study types may help to elucidate the directionality of relationships between variables at a very individual level by modelling the complex interdependencies between multiple dimensions (e.g., genome, transcriptome, epigenome, microbiome, and metabolome), thereby moving away from associative to a more causal analysis. To ascertain the role of machine learning in allergy research, we conducted a comprehensive systematic review of the current literature. The findings highlight and underscore the potential of using AI/ML approaches in advancing our understanding of allergic diseases, which ultimately enhances patient care through improved prevention, diagnosis, and management strategies. It is important to emphasise that there is no single 'best' analytical method, highlighting the importance of cross-disciplinary collaborations. A team science approach is crucial for ensuring the application of appropriate methodologies tailored to the research question at hand and that context-specific interpretations are being made, supported by critical appraisal from both the front- (e.g., clinicians) and back-end (e.g., analysts) of research processes.
BACKGROUND:Component-resolved diagnostics allow detailed assessment of IgE sensitization to multiple allergenic molecules (component-specific IgEs, or c-sIgEs) and may be useful for asthma diagnosis. However, to effectively use component-resolved diagnostics across diverse settings, it is crucial to account for geographic differences. OBJECTIVE:We investigated spatial determinants of c-sIgE networks to facilitate development of diagnostic algorithms applicable globally. METHODS:We used multiplex component-resolved diagnostics array to measure c-sIgE to 112 proteins in an international collaboration of several studies: WASP (World Asthma Phenotypes; United Kingdom, New Zealand, Brazil, Ecuador, and Uganda), U-BIOPRED (Unbiased Biomarkers for the Prediction of Respiratory Disease Outcomes; 7 European countries), and MAAS (Manchester Asthma and Allergy Study, a UK population-based birth cohort). Hierarchical clustering on low-dimensional representation of co-occurrence networks ascertained sensitization and c-sigE clusters across populations. Cross-country comparisons focused on a common subset of 18 c-sIgEs. We investigated sensitization networks across regions in relation to asthma severity. RESULTS:Sensitization profiles shared similarities across regions. For 18 c-sIgEs shared across study populations, the response structure enabled differentiation between different geographic areas and study designs, revealing 3 clusters: (1) Uganda, Ecuador, and Brazil, (2) U-BIOPRED children and adults, and (3) New Zealand, United Kingdom, and MAAS. Spectral clustering identified differences between clusters. We observed constant, almost parallel shifts between severe and nonsevere asthma in each country. CONCLUSIONS:Patterns of c-sIgE response reflect geographic location and study design. However, despite geographic differences in c-sIgE networks, there is a remarkably consistent shift between networks of subjects with nonsevere and severe asthma.
Textual data analysis is critical for monitoring changing themes over time. To overcome challenges posed by data richness, graph theory emerges as a tool for investigating word-topic associations. We present an approach to clustering co-occurrence word networks that prioritises network similarity quantification over time. Addressing theoretical and network geometrical constraints, a statistical framework for manifold data analysis facilitates the grouping of semantic networks, partitioning the observed time frame into periods, and identifying dominant topics in each period via tensor decomposition. The analysis of Brexit-related tweets demonstrates the efficacy of modern methods for identifying social media patterns on public discourse.
INTRODUCTION: Asthma is a chronic respiratory disease that affects millions of people worldwide, and despite intensive study, the underlying molecular pathways are still unknown. Recent breakthroughs in machine learning and multi-omics technology provide new avenues for delving deeper into disease pathophysiology and identifying potential treatment targets. EVIDENCE ACQUISITION: We comprehensively reviewed the literature to explore the potential of machine learning and multi-omics approaches in asthma research. We searched the Scopus database using a combination of terms such as "asthma," "machine learning", and "multi-omics" and their synonyms. EVIDENCE SYNTHESIS: Our review revealed that machine learning and multi-omics approaches have been increas- ingly used to identify biomarkers, classify asthma subtypes, predict treatment outcomes, and understand the disease pathophysiology at the molecular level. CONCLUSIONS: Combining machine learning and multi-omics technologies holds tremendous potential for advancing our understanding of asthma pathogenesis and identifying novel therapeutic targets. Overall, this analysis demonstrates how these techniques have the potential to revitalize asthma research and improve patient outcomes. More study is needed, however, to confirm the utility of these approaches and understand how to build viable models that can be applied to clinical practice. (Cite this article as: Fontanella S, Cucco A, Custovic A. Breathing new life into asthma research: a review of machine learning and multi-omics approaches. Minerva Respir Med 2023;62:163-76. DOI: 10.23736/S2784-8477.23.02068-5)
Introduction: Big data are reshaping the future of medicine. The growing availability and increasing complexity of data have favored the adoption of modern analytical and computational methodologies in every area of medicine. Over the past decades, asthma research has been characterized by a shift in the way studies are conducted and data are analyzed. Motivated by the assumptions that 'data will speak for themselves', hypothesis-driven approaches have been replaced by data-driven hypotheses-generating methods to explore hidden patterns and underlying mechanisms. However, even with all the advancement in technologies and the new important insight that we gained to understand and characterize asthma heterogeneity, very few research findings have been translated into clinically actionable solutions. Areas covered: To investigate some of the fundamental analytical approaches adopted in the current literature and appraise their impact and usefulness in medicine, we conducted a bibliometric analysis of big data analytics in asthma research in the past 50 years. Expert opinion: No single data source or methodology can uncover the complexity of human health and disease. To fully capitalize on the potential of 'big data', we will have to embrace the collaborative science and encourage the creation of integrated cross-disciplinary teams brought together around technological advances.