BACKGROUND:Patients with congenital heart disease are identified in 1% of live births. Improved surgical intervention means many patients now survive to adulthood, the corollary of which is increased mortality in the over-65-year-old congenital heart disease (CHD) population. In the clinic, genetic sequencing increasingly identifies novel genetic variants in genes related to CHD. Traditional assays for interpreting novel genetic variants are often limited by gene-specificity, whereas animal models are cumbersome and may not accurately reflect human disease. This study investigates CRISPR gene editing in induced pluripotent stem cells and cardiomyocyte-directed differentiation as a human disease model to investigate novel genetic variants identified in association with CHD. METHODS AND RESULTS:We identified a GATA4 p.Arg284His genetic variant in a paediatric patient. This genetic variant was introduced into induced pluripotent stem cells (iPSCs) using CRISPR gene editing with homology-directed-repair. GATA4 genetic variant and isogenic control iPSCs were selected and differentiated into cardiomyocytes. Expression of the GATA4 p.Arg284His variant resulted in altered calcium transients, indicative of CHD and consistent with the patient's clinical phenotype. Transcriptomics revealed cellular pathway changes in cardiac development, calcium handling, and energy metabolism that contribute to disease aetiology, mechanism and identification of potential treatments. CONCLUSION:Directed differentiation of iPSCs harbouring the GATA4 p.Arg284His genetic variant recapitulated the CHD phenotype, indicated disease mechanisms, and pointed to potential sites for targeting with therapy. The study highlights the utility of transcriptomics for the functional interpretation of cardiac genetic variants and is an exemplar for precision medicine approaches for the investigation of CHD.
An artificial-intelligence system uses clinical data, genetic information and literature searches to suggest diagnoses and provides the underlying reasoning. An artificial-intelligence system uses clinical data, genetic information and literature searches to suggest diagnoses and provides the underlying reasoning.
Gene set analysis often returns extensive annotations from multiple sources, requiring manual effort to identify coherent biological themes. We developed GeneInsight, an AI-powered tool that automates this by retrieving functional annotations from STRING-DB, clustering semantically related terms using sentence embeddings, and generating thematic summaries via large language model prompting. This enables researchers to identify biological themes that may be obscured when annotation sources are examined separately.
Background: An estimated 1 in 12 individuals across the world have a rare disease. The Undiagnosed Diseases Network International (UDNI) recommend Undiagnosed Disease Programs (UDPs) as the best approach to facilitate diagnosis and support for those affected. The Australian Undiagnosed Disease Network (UDN-Aus) is the first National Australian UDP initiative funded by the Medical Research Future Fund’s Genomic Health Futures Mission (GNT2007567). Aims: UDN-Aus brings together an unprecedented national collaborative network for rare disease, aiming to improve the rate of genomic diagnoses for those with undiagnosed rare genetic conditions, enabling precise, personalised care to individuals throughout Australia. Methods: The research program recruited Australians who have been seen through a clinical genetics service and remained undiagnosed following clinically available genomic testing. This paper outlines the approach taken to establish a national research project at 12 clinical recruitment sites. The methods detail the funding, aims, governance, study design, population and participation process, health economic research and preliminary results, and recommendations for future sustainability and implementation. Results: The study was approved by the Royal Children’s Hospital Human Research Ethics Committee on 19 November 2021 (RCH79712), with relevant site-specific approvals at local recruitment sites. Key benefits and barriers in the establishment of UDN-Aus are outlined. Conclusion: The successful establishment of this program required several components, including meaningful and ongoing community and stakeholder engagement, strategic appointment of key operational staff, and a tailored approach to facilitate more equitable enrolment. It also highlighted several imperative areas for consideration to ensure future sustainable implementation of a national UDP. These include continued investment in Australia’s national genomic data transfer policy and infrastructure, and research ethics and governance procedures is imperative for the sustainable delivery of genomic research for rare disease.
Mesothelioma is a cancer derived from mesothelial cells, most commonly arising from the pleura or the peritoneum. Immune checkpoint therapy (ICT) has shown survival benefit for pleural mesothelioma, but little is known about the response in peritoneal mesothelioma. Most preclinical mesothelioma models involve subcutaneous cancer cell implantation, which lacks the relevant tumour microenvironment of peritoneal mesothelioma and does not resemble the clinical presentation. We therefore set out to explore the influence of location on the mesothelioma tumour microenvironment, comparing pleural, peritoneal, and subcutaneous models using identical cell line-derived syngeneic mesotheliomas. We found that the peritoneal location conferred an anti-inflammatory tumour microenvironment, characterised by low IFN, TNFα, and STAT signalling activity, low immune cell infiltration and a gene signature associated with non-response to ICT. ICT was effective in subcutaneous models, but the same cell line-derived tumours were irresponsive when inoculated intraperitoneally. Together, these findings show that peritoneal location is associated with an immune suppressive tumour microenvironment.
Large amounts of transcriptomic data have been made available in public repositories. Systematic reanalyses of these data offer the potential to identifying conserved biological patterns or context-specific signatures. However, this is a labour intensive process requiring bioinformatic expertise and a long chain of manual decision making. Use of LLMs and agentic systems holds promise for automating these otherwise time-consuming tasks. Here, we present UORCA (Unified -Omics Reference Corpus of Analyses), a tool to systematically identify and analyse public transcriptomic datasets. UORCA uses an LLM-assisted framework to search for datasets relevant to a research question. These datasets are analysed through a multi-agent system that performs a standardised bioinformatic analyses to identify differentially expressed genes. Results of each analysis are then displayed in an interactive visual interface. We found that UORCA recapitulated findings reported from a manual comparison of datasets, but also found biological signatures that were not initially described. We find that UORCA generates targeted hypotheses relevant for drug design, and facilitates evaluation of experimental results where they differ from past literature. Together, these findings demonstrate how UORCA accelerates biomedical discovery by enabling scientists to extract actionable findings from diverse public datasets.
Abstract The developing immune system of a child is distinct to that of an adult. These immunological differences are often ignored in preclinical pediatric cancer research, where adult mice are more commonly used, potentially overlooking developmental influences on cancer-microenvironment interactions. This is of particular importance when testing immunotherapeutic agents for pediatric cancers. To address this issue, we have developed pediatric brain cancer mouse models which reflect the developing microenvironment in which these tumors arise. By doing so, we sought to understand the impact of age on tumor progression and immune interactions. Using flow cytometry, RNA sequencing, and immunohistochemistry, we have characterized differences in the tumor-immune microenvironment of multiple orthotopically-implanted murine brain tumor models in juvenile mice compared to adults. We found that identical brain tumor cells elicited tumors that grew faster in juvenile mice and had fewer immune cell infiltrates. Moreover, these immune infiltrates were markedly distinct between juvenile and adult mice. Specifically, juvenile mice possessed more naïve-like CD8 T cells with reduced effector, resident, and exhausted-like CD8 T cells. Tumor-associated macrophages in juvenile mice had reduced MHC II expression and appeared polarized towards an anti-inflammatory state, potentially suppressing effective anti-tumour immune responses. Importantly, we demonstrate that repolarization of macrophages using immune-modulating agents changed the pediatric tumor-infiltrating immune microenvironment towards a more “adult-like state”, that may enhance immunotherapy effectiveness. Acknowledging the challenges in finding an appropriate match for human developmental stage in mice, our findings highlight that preclinical model age significantly influences cancer-immune interactions. These data strongly support the use of age-relevant models in preclinical pediatric cancer studies, especially when evaluating microenvironment-targeting agents.
Seven female individuals with multiple congenital anomalies, developmental delay and/or intellectual disability have been found to have a genetic variant of uncertain significance in the mediator complex subunit 12 gene (MED12 c.3412C>T, p.Arg1138Trp). The functional consequence of this genetic variant in disease is undetermined, and insight into disease mechanism is required. We identified a de novo MED12 p.Arg1138Trp variant in a female patient and compared disease phenotypes with six female individuals identified in the literature. To investigate affected biological pathways, we derived two induced pluripotent stem cell (iPSC) lines from the patient: one expressing wildtype MED12 and the other expressing the MED12 p.Arg1138Trp variant. We performed neural disease modelling, transcriptomics and protein analysis, comparing healthy and variant cells. When comparing the two cell lines, we identified altered gene expression in neural cells expressing the variant, including genes regulating RNA polymerase II activity, transcription, pre-mRNA processing, and neural development. We also noted a decrease in MED12L expression. Pathway analysis indicated temporal delays in axon development, forebrain differentiation, and neural cell specification with significant upregulation of pre-ribosome complex gene pathways. In a human neural model, expression of MED12 p.Arg1138Trp altered neural cell development and dysregulated the pre-ribosome complex providing functional evidence of disease aetiology and mechanism in MED12-related disorders.
Interpreting gene sets is often complicated by the overwhelming number of annotations associated with individual genes, making it difficult to extract meaningful biological insights. To address this issue, we developed GeneInsight, an AI-powered tool that combines advanced topic modelling with large language models to automatically synthesise diverse biological annotations from literature, gene ontologies, and databases such as STRING. GeneInsight consolidates extensive annotations into coherent thematic summaries that render such data readily interpretable, thereby enabling the rapid extraction of biologically significant insights that conventional enrichment analyses often overlook. ### Competing Interest Statement The authors have declared no competing interest.
Immune checkpoint therapy (ICT) causes durable tumour responses in a subgroup of patients, but it is not well known how T cell receptor beta (TCRβ) repertoire dynamics contribute to the therapeutic response. Using murine models that exclude variation in host genetics, environmental factors and tumour mutation burden, limiting variation between animals to naturally diverse TCRβ repertoires, we applied TCRseq, single cell RNAseq and flow cytometry to study TCRβ repertoire dynamics in ICT responders and non-responders. Increased oligoclonal expansion of TCRβ clonotypes was observed in responding tumours. Machine learning identified TCRβ CDR3 signatures unique to each tumour model, and signatures associated with ICT response at various timepoints before or during ICT. Clonally expanded CD8+ T cells in responding tumours post ICT displayed effector T cell gene signatures and phenotype. An early burst of clonal expansion during ICT is associated with response, and we report unique dynamics in TCRβ signatures associated with ICT response.
Time-critical transcriptional events in the immune microenvironment are important for response to immune checkpoint blockade (ICB), yet these events are difficult to characterise and remain incompletely understood. Here, we present whole tumor RNA sequencing data in the context of treatment with ICB in murine models of AB1 mesothelioma and Renca renal cell cancer. We sequenced 144 bulk RNAseq samples from these two cancer types across 4 time points prior and after treatment with ICB. We also performed single-cell sequencing on 12 samples of AB1 and Renca tumors an hour before ICB administration. Our samples were equally distributed between responders and non-responders to treatment. Additionally, we sequenced AB1-HA mesothelioma tumors treated with two sample dissociation protocols to assess the impact of these protocols on the quality transcriptional information in our samples. These datasets provide time-course information to transcriptionally characterize the ICB response and provide detailed information at the single-cell level of the early tumor microenvironment prior to ICB therapy.
A robust understanding of the cellular mechanisms underlying diseases sets the foundation for the effective design of drugs and other interventions. The wealth of existing single-cell atlases offers the opportunity to uncover high-resolution information on expression patterns across various cell types and time points. To better understand the associations between cell types and diseases, we leveraged previously developed tools to construct a standardized analysis pipeline and systematically explored associations across four single-cell datasets, spanning a range of tissue types, cell types and developmental time periods. We utilized a set of existing tools to identify co-expression modules and temporal patterns per cell type and then investigated these modules for known disease and phenotype enrichments. Our pipeline reveals known and novel putative cell type-disease associations across all investigated datasets. In addition, we found that automatically discovered gene co-expression modules and temporal clusters are enriched for drug targets, suggesting that our analysis could be used to identify novel therapeutic targets.
Genetic diagnosis plays a crucial role in rare diseases, particularly with the increasing availability of emerging and accessible treatments. The International Rare Diseases Research Consortium (IRDiRC) has set its primary goal as: “Ensuring that all patients who present with a suspected rare disease receive a diagnosis within one year if their disorder is documented in the medical literature”. Despite significant advances in genomic sequencing technologies, more than half of the patients with suspected Mendelian disorders remain undiagnosed. In response, IRDiRC proposes the establishment of “a globally coordinated diagnostic and research pipeline”. To help facilitate this, IRDiRC formed the Task Force on Integrating New Technologies for Rare Disease Diagnosis. This multi-stakeholder Task Force aims to provide an overview of the current state of innovative diagnostic technologies for clinicians and researchers, focusing on the patient’s diagnostic journey. Herein, we provide an overview of a broad spectrum of emerging diagnostic technologies involving genomics, epigenomics and multi-omics, functional testing and model systems, data sharing, bioinformatics, and Artificial Intelligence (AI), highlighting their advantages, limitations, and the current state of clinical adaption. We provide expert recommendations outlining the stepwise application of these innovative technologies in the diagnostic pathways while considering global differences in accessibility. The importance of FAIR (Findability, Accessibility, Interoperability, and Reusability) and CARE (Collective benefit, Authority to control, Responsibility, and Ethics) data management is emphasized, along with the need for enhanced and continuing education in medical genomics. We provide a perspective on future technological developments in genome diagnostics and their integration into clinical practice. Lastly, we summarize the challenges related to genomic diversity and accessibility, highlighting the significance of innovative diagnostic technologies, global collaboration, and equitable access to diagnosis and treatment for people living with rare disease.
BackgroundSETBP1 Haploinsufficiency Disorder (SETBP1-HD) is characterised by mild to moderate intellectual disability, speech and language impairment, mild motor developmental delay, behavioural issues, hypotonia, mild facial dysmorphisms, and vision impairment. Despite a clear link between SETBP1 mutations and neurodevelopmental disorders the precise role of SETBP1 in neural development remains elusive. We investigate the functional effects of three SETBP1 genetic variants including two pathogenic mutations p.Glu545Ter and SETBP1 p.Tyr1066Ter, resulting in removal of SKI and/or SET domains, and a point mutation p.Thr1387Met in the SET domain.MethodsGenetic variants were introduced into induced pluripotent stem cells (iPSCs) and subsequently differentiated into neurons to model the disease. We measured changes in cellular differentiation, SETBP1 protein localisation, and gene expression changes.ResultsThe data indicated a change in the WNT pathway, RNA polymerase II pathway and identified GATA2 as a central transcription factor in disease perturbation. In addition, the genetic variants altered the expression of gene sets related to neural forebrain development matching characteristics typical of the SETBP1-HD phenotype.LimitationsThe study investigates changes in cellular function in differentiation of iPSC to neural progenitor cells as a human model of SETBP1 HD disorder. Future studies may provide additional information relevant to disease on further neural cell specification, to derive mature neurons, neural forebrain cells, or brain organoids.ConclusionsWe developed a human SETBP1-HD model and identified perturbations to the WNT and POL2RA pathway, genes regulated by GATA2. Strikingly neural cells for both the SETBP1 truncation mutations and the single nucleotide variant displayed a SETBP1-HD-like phenotype.
In the first-ever Undiagnosed Hackathon, nearly 100 experts from 28 countries combined advanced phenotyping and genomic techniques for 48 hours, ultimately providing diagnoses to 40% of the previously undiagnosed families. This inspiring model demonstrates the power of multidisciplinary collaboration and patient partnership in precision diagnostics.
MOTIVATION:Over the last two decades, transcriptomics has become a standard technique in biomedical research. We now have large databases of RNA-seq data, accompanied by valuable metadata detailing scientific objectives and the experimental procedures used. The metadata is crucial in understanding and replicating published studies, but so far has been underutilized in helping researchers to discover existing datasets. RESULTS:We present SampleExplorer, a tool allowing researchers to search for relevant data using both text and gene set queries. SampleExplorer embeds sample metadata and uses a transformer-based language model to retrieve similar datasets. Extensive benchmarking (see Supplementary Materials and Methods) using the ARCHS4 database demonstrates that SampleExplorer provides an effective approach for retrieving biologically relevant samples from large-scale transcriptomicdata. This tool provides an efficient approach for discovering relevant gene expression datasets in large public repositories. It improves sample and dataset identification across diverse experimental contexts, helping researchers leverage existing transcriptomic data for potential replication or verification studies. Availability and implementation: SampleExplorer is available as a Python package compatible with versions 3.9 to 3.11, available for installation via the Python Package Index (PyPI). The codebase and documentation are accessible at https://github.com/wlchin/SampleExplorer. Supplementary data (Supplementary Materials and Methods) provides detailed methodological information, including an algorithmic description of the retrieval process and data preparation steps.