BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., ‘CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W’ that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (relative to that event); and (ii) a compact version named ‘stub’ (e.g., ‘CT01001LTR0N401T1W’) optimized for filenames, pipelines, and labeling. ClarID is supported by an open-source reference implementation, ClarID-Tools, a command-line tool that processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID’s utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID is designed to complement, not replace, persistent identifiers such as UUIDs, by providing a human-readable layer that enhances interpretability and facilitates metadata curation. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools .
Motivation:Variant calling for next-generation sequencing (NGS) data relies on a diverse ecosystem of tools and workflows. Large-scale collaborative studies increasingly adopt federated analysis, where each institution processes sensitive data locally using standardized pipelines. Deploying identical pipelines across multiple centers remains challenging because heterogeneous software environments and computing policies can cause workflow divergence and inconsistent results. Results:We developed CBIcall, a workflow backend-flexible, configuration-driven framework that runs standardized variant-calling pipelines from raw FASTQ files to analysis-ready VCFs. Users define each analysis in a single YAML parameters file, which CBIcall resolves against a controlled workflow registry and resource catalog. The execution driver validates parameters and checks compatibility among pipelines, analysis modes, workflow backends, genome builds, tool versions, and resource bundles. CBIcall supports reproducibility auditing by comparing executions using recorded provenance and output fingerprints. CBIcall dispatches validated workflows natively through Bash, Cromwell, Nextflow and Snakemake backends and provides production-ready pipelines for germline WES, WGS (single-sample or cohort joint genotyping following GATK Best Practices), and mitochondrial DNA analysis. We evaluated analytical performance using public benchmark datasets and validated reproducibility across four computing environments. We further deployed CBIcall in the EU HEREDITARY project, where it processed 1102 samples with both WES and mtDNA pipelines on an institutional HPC system, supporting its suitability for reproducible cohort-scale genomic analyses. Availability and implementation:CBIcall is open source (GPLv3) and distributed with ready-to-run pipelines; full dependency and installation documentation is available at https://github.com/CNAG-Biomedical-Informatics/cbicall.
Our immune system is constantly exposed to fungi, but it mounts a consistent, disease-related, and species-specific inflammatory response only against a few fungal species. Most of the current understanding of fungus-host interactions is based on a limited number of strains, hence neglecting the fungal intra-specific genetic and phenotypic diversity. To expand our knowledge of the spectrum of immune responses to pathogenic and non-pathogenic fungi, we compared the cytokine and transcriptional profiles of human monocyte-derived dendritic cells exposed to Aspergillus fumigatus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae strains. The tested species triggered common and species-specific responses, mostly resulting from the different timing of signaling pathways. Faster phagolysosome acidification was observed for pathogenic species. These results highlight the urgency to redraw the boundaries between pathogenicity and commensalism in fungi, shining a spotlight on the timing of the response, rather than solely on the genes triggered by the stimuli.
Most populations of spider monkeys ( Ateles ) and muriquis ( Brachyteles ), two Neotropical primate genera, are under severe anthropogenic threats. Yet, taxon-wide population-level studies leveraging their degree of endangerment linked to their genetic diversity patterns and demographic history are lacking. To properly address this, there is a need to expand from morphological and genetic marker-based studies. We generated high-coverage genome sequencing for 58 individuals sampled across 8 Atelidae species, in the first population-wide study of all extant spider monkey species, in the wild and captivity, alongside northern muriquis ( Brachyteles hypoxanthus ). Additionally, we present a high-contiguity reference genome for Ateles hybridus . Here, we observe the overall levels of genetic diversity and genetic burden of the analyzed populations do not align to their IUCN endangerment category. Moreover, we show that in the wild, genetic burden is overall higher compared to the captive populations analyzed. Then, we depict two main trans and cis-Andean sister clades in Ateles , and further structure and dynamics outlined by the Madeira River in the latter clade. Lastly, we find that genes in highly divergent regions between Ateles and B. hypoxanthus are involved in central nervous system development and photorreception. Our study shows i) the lack of concordance between the genetic diversity levels and extinction risk of these populations, suggestive of recent and strong external drivers; ii) increased genetic burden in the wild in contrast to effective captive management, indicating mostly past demographic events; iii) structure and dynamics in spider monkeys that agrees with common biogeographical patterns and iv) genetic divergence between Ateles and Brachyteles potentially linked to distinct environmental light levels. ### Competing Interest Statement Employees of Illumina, Inc. are indicated in the list of author affiliations. The other authors declare no competing interests.
Long-read sequencing technologies enable resolution of structural variants (SV) and long-range genome assembly, but require high molecular weight (HMW) DNA of both high quantity and quality to produce optimal sequencing results. New DNA extraction methods have been developed but these have not been assessed for use in routine testing. The interlaboratory study described here tested four commonly used methods: Fire Monkey, Nanobind, Puregene and Genomic-tip with a reference cell line containing known chromosomal alterations. Samples were assessed with commonly applied approaches for evaluating DNA purity and integrity as well as a method based on linkage using digital PCR. Sequencing performance was evaluated and the impact of extraction method on structural variant calling investigated. All methods generally produced samples of acceptable purity although yield varied considerably between laboratories. Library preparation and sequencing were successful for all four methods, with Fire Monkey extracts achieving the highest N50 values, Genomic Tip giving the highest sequencing yields and Nanobind, the highest proportion of ultra-long reads (> 100 kb). The dPCR assay with duplexes at 100 kb and 150 kb distances was predictive of ultra-long reads and provides a more quantitative read-out (
Psoroptes ovis is a mite species that feeds on sheep, cattle, other ungulates, rabbits, and horses, which can develop into a severe exudative dermatitis known as psoroptic mange. The macrocyclic lactone (ML) family of acaricides are commonly used to control psoroptic mange. However, certain strains of cattle and sheep mites have developed resistance against MLs, which has led to reduced treatment efficacy and even treatment failure. Here we investigated the genetic basis of ML resistance in P. ovis mites collected from cattle across Belgium. We compared gene expression between susceptible and resistant mites before and after exposure to ivermectin and genetic diversity between a single susceptible and resistant populations. We generated chromosomal genome assemblies of P. ovis derived from sheep and cattle respectively and correlated genomic diversity of susceptible and resistant P. ovis populations sampled across Belgium. Gene expression data revealed constitutive over-expression of a cytochrome P450 monooxygenase (CYP) gene and two tandemly located UDP-glucuronosyltransferase (UGT) genes among others. On investigation of the genomic data, we observed copy number variation at both loci in population genomic data. The CYP gene is not amplified in the susceptible population but occurs in multiple copies in all resistant populations and is associated with a peak in Fst between resistant and susceptible populations indicative of selection. By contrast, the two UGT genes are massively and tandemly amplified in all populations including the susceptible dataset with weaker Fst between populations than the amplified CYP gene. Hence, distinct mechanisms of amplification and gene regulation are occurring at these putative resistance loci in P. ovis.
Cancer development and response to treatment are evolutionary processes1,2, but characterizing evolutionary dynamics at a clinically meaningful scale has remained challenging3. Here we develop a new methodology called EVOFLUx, based on natural DNA methylation barcodes fluctuating over time4, that quantitatively infers evolutionary dynamics using only a bulk tumour methylation profile as input. We apply EVOFLUx to 1,976 well-characterized lymphoid cancer samples spanning a broad spectrum of diseases and show that initial tumour growth rate, malignancy age and epimutation rates vary by orders of magnitude across disease types. We measure that subclonal selection occurs only infrequently within bulk samples and detect occasional examples of multiple independent primary tumours. Clinically, we observe faster initial tumour growth in more aggressive disease subtypes, and that evolutionary histories are strong independent prognostic factors in two series of chronic lymphocytic leukaemia. Using EVOFLUx for phylogenetic analyses of aggressive Richter-transformed chronic lymphocytic leukaemia samples detected that the seed of the transformed clone existed decades before presentation. Orthogonal verification of EVOFLUx inferences is provided using additional genetic data, including long-read nanopore sequencing, and clinical variables. Collectively, we show how widely available, low-cost bulk DNA methylation data precisely measure cancer evolutionary dynamics, and provides new insights into cancer biology and clinical behaviour.
Chronic lymphocytic leukemia is a complex and heterogeneous hematological malignancy. The advance of high-throughput multi-omics technologies has significantly influenced chronic lymphocytic leukemia research and paved the way for precision medicine approaches. In this review, we explore the role of machine learning in the analysis of multi-omics data in this hematological malignancy. We discuss recent literature on different machine learning models applied to single omic studies in chronic lymphocytic leukemia, with a special focus on the potential contributions to precision medicine. Finally, we highlight the recently published machine learning applications in multi-omics data in this area of research as well as their potential and limitations.
AbstractPolymyositis with mitochondrial pathology (PM-Mito) was first identified in 1997 as a subtype of idiopathic inflammatory myopathy. Recent findings demonstrated significant molecular similarities between PM-Mito and Inclusion Body Myositis (IBM), suggesting a trajectory from early to late IBM and prompting the inclusion of PM-Mito as an IBM precursor (early IBM) within the IBM spectrum. Both PM-Mito and IBM show mitochondrial abnormalities, suggesting mitochondrial disturbance is a critical element of IBM pathogenesis.The primary objective of this cross-sectional study was to characterize the mitochondrial phenotype in PM-Mito at histological, ultrastructural, and molecular levels and to study the interplay between mitochondrial dysfunction and inflammation. Skeletal muscle biopsies of 27 patients with PM-Mito and 27 with typical IBM were included for morphological and ultrastructural analysis. Mitochondrial DNA (mtDNA) copy number and deletions were assessed by qPCR and long-range PCR, respectively. In addition, full-length single-molecule sequencing of the mtDNA enabled precise mapping of deletions. Protein and RNA levels were studied using unbiased proteomic profiling, immunoblotting, and bulk RNA sequencing. Cell-free mtDNA (cfmtDNA) was measured in the serum of IBM patients.We found widespread mitochondrial abnormalities in both PM-Mito and IBM, illustrated by elevated numbers of COX-negative and SDH-positive fibers and prominent ultrastructural abnormalities with disorganized and concentric cristae within enlarged and dysmorphic mitochondria. MtDNA copy numbers were significantly reduced, and multiple large-scale mtDNA deletions were already evident in PM-Mito, compared to healthy age-matched controls, similar to the IBM group. The activation of the canonical cGAS/STING inflammatory pathway, possibly triggered by the intracellular leakage of mitochondrial DNA, was evident in PM-Mito and IBM. Elevated levels of circulating cfmtDNA also indicated leakage of mtDNA as a likely inflammatory trigger. In PM-Mito and IBM, these findings were accompanied by dysregulation of proteins and transcripts linked to the mitochondrial membranes.In summary, we identified that mitochondrial dysfunction with multiple mtDNA deletions and depletion, disturbed mitochondrial ultrastructure, and defects of the inner mitochondrial membrane are features of PM-Mito and IBM, underlining the concept of an IBM-spectrum disease (IBM-SD). The activation of inflammatory pathways related to mtDNA release indicates a significant role of mitochondria-associated inflammation in the pathogenesis of IBM-SD. Thus, mitochondrial abnormalities precede tissue remodeling and infiltration by specific T-cell subpopulations (e.g., KLRG1+) characteristic of late IBM. This study highlights the critical role of early mitochondrial abnormalities in the pathomechanism of IBM, which may lead to new approaches to therapy.
Despite showing the greatest primate diversity on the planet, genomic studies on Amazonian primates show very little representation in the literature. With 48 geolocalized high coverage whole genomes from wild uakari monkeys, we present the first population-level study on platyrrhines using whole genome data. In a very restricted range of the Amazon rainforest, eight uakari species (Cacajao genus) have been described and categorized into the bald and black uakari groups, based on phenotypic and ecological differences. Despite a slight habitat overlap, we show that posterior to their split 0.92 Mya, bald and black uakaris have remained independent, without gene flow. Nowadays, these two groups present distinct genetic diversity and group-specific variation linked to pathogens. We propose differing hydrology patterns and effectiveness of geographic barriers have modulated the intra-group connectivity and structure of bald and black uakari populations. With this work we have explored the effects of the Amazon rainforest’s dynamism on wild primates’ genetics and increased the representation of platyrrhine genomes, thus opening the door to future research on the complexity and diversity of primate genomics. Population study of whole genomes of wild uakary monkeys (Cacajao genus) exposes how the dynamic and highly heterogeneous Amazon basin may have shaped complex connectivity patterns and driven fast population differentiation on these.
The use of single-cell technologies for clinical applications requires disconnecting sampling from downstream processing steps. Early sample preservation can further increase robustness and reproducibility by avoiding artifacts introduced during specimen handling. We present FixNCut, a methodology for the reversible fixation of tissue followed by dissociation that overcomes current limitations. We applied FixNCut to human and mouse tissues to demonstrate the preservation of RNA integrity, sequencing library complexity, and cellular composition, while diminishing stress-related artifacts. Besides single-cell RNA sequencing, FixNCut is compatible with multiple single-cell and spatial technologies, making it a versatile tool for robust and flexible study designs.
The Catalan Initiative for the Earth BioGenome Project (CBP) is an EBP-affiliated project network aimed at sequencing the genome of the >40 000 eukaryotic species estimated to live in the Catalan-speaking territories (Catalan Linguistic Area, CLA). These territories represent a biodiversity hotspot. While covering less than 1% of Europe, they are home to about one fourth of all known European eukaryotic species. These include a high proportion of endemisms, many of which are threatened. This trend is likely to get worse as the effects of global change are expected to be particularly severe across the Mediterranean Basin, particularly in freshwater ecosystems and mountain areas. Following the EBP model, the CBP is a networked organization that has been able to engage many scientific and non-scientific partners. In the pilot phase, the genomes of 52 species are being sequenced. As a case study in biodiversity conservation, we highlight the genome of the Balearic shearwater Puffinus mauretanicus, sequenced under the CBP umbrella.
Palatine tonsils are secondary lymphoid organs (SLOs) representing the first line of immunological defense against inhaled or ingested pathogens. We generated an atlas of the human tonsil composed of >556,000 cells profiled across five different data modalities, including single -cell transcriptome, epigenome, proteome, and immune repertoire sequencing, as well as spatial transcriptomics. This census identified 121 cell types and states, defined developmental trajectories, and enabled an understanding of the functional units of the tonsil. Exemplarily, we stratified myeloid slan-like subtypes, established a BCL6 enhancer as locally active in follicle -associated T and B cells, and identified SIX5 as putative transcriptional regulator of plasma cell maturation. Analyses of a validation cohort confirmed the presence, annotation, and markers of tonsillar cell types and provided evidence of age -related compositional shifts. We demonstrate the value of this resource by annotating cells from B cell -derived mantle cell lymphomas, linking transcriptional heterogeneity to normal B cell differentiation states of the human tonsil.
Efficient sharing and integration of phenotypic data is crucial for advancing biomedical research and enhancing patient outcomes in precision medicine and public health. To achieve this, the health data community has developed standards to promote the harmonization of variable names and values. However, the use of diverse standards across different research centers can hinder progress. Here we present Convert-Pheno, an open-source software toolkit that enables the interconversion of common data models for phenotypic data such as Beacon v2 Models, CDISC-ODM, OMOP-CDM, Phenopackets v2, and REDCap. Along with the software, we have created a detailed documentation that includes information on deployment and installation.
Abstract Genomics data will soon be routinely generated and integrated into national healthcare systems. To maximise the potential of genomic medicine, data should be accessible for research where possible. This is the remit of ELIXIR, an intergovernmental organisation that brings together life science resources from across Europe. Innovative solutions are needed to ensure the validated research findings for disease or preventative medicine are then integrated into healthcare. To tackle this, the 1+MG initiative, a joint initiative of 25 EU countries, the UK, and Norway, aims to enable secure access to genomics and the corresponding clinical data across Europe for better research, personalised healthcare and health policy making. In the design and scale-up phase (B1MG project), recommendations and guidelines to advance towards the deployment of personalised medicine at a European scale have been produced, adopted by 1+MG and developed into a 1+MG framework. This includes guidance on data governance, standards, quality and infrastructure, recommendations on how to approach citizen engagement and a tool for countries to self-assess implementation into healthcare. The European Genomic Data Infrastructure (GDI) project supports the scale-up and sustainability phase of the 1+MG initiative, to deploy infrastructure across 24 countries to support the overall ambition. Recommendations are being used to promote governance and technical interoperability across European initiatives including the European Health Data Space (EHDS), and European Cancer Image Initiative (EUCAIM). The 1+MG will be established as a European Data Infrastructure Consortia in 2025 and will act as an Authorised Participant in the EHDS providing access to a permanent high-quality federated data collection of genomic and health data that will accelerate research, innovation and policymaking facilitating the deployment of genomic medicine across Europe.
Ecological variation and anthropogenic landscape modification have had key roles in the diversification and extinction of mammals in Madagascar. Lemurs represent a radiation with more than 100 species, constituting roughly one-fifth of the primate order. Almost all species of lemurs are threatened with extinction, but little is known about their genetic diversity and demographic history. Here, we analyse high-coverage genome-wide resequencing data from 162 unique individuals comprising 50 species of Lemuriformes, including multiple individuals from most species. Genomic diversity varies widely across the infraorder and yet is broadly consistent among individuals within species. We show widespread introgression in multiple genera and generally high levels of genomic diversity likely resulting from allele sharing that occurred during periods of connectivity and fragmentation during climatic shifts. We find distinct patterns of demographic history in lemurs across the ecogeographic regions of Madagascar within the last million years. Within the past 2,000 years, lemurs underwent major declines in effective population size that corresponded to the timing of human population expansion in Madagascar. In multiple regions of the island, we identified chronological trajectories of inbreeding that are consistent across genera and species, suggesting localized effects of human activity. Our results show how the extraordinary diversity of these long-neglected, endangered primates has been influenced by ecological and anthropogenic factors. Analysis of the genomes of 50 species of Lemuriformes shows high levels of genomic diversity, likely due to allele sharing, as well as population declines and inbreeding patterns resulting from ecological factors and human impacts in Madagascar.
Despite the wealth of publicly available single-cell datasets, our understanding of distinct resident immune cells and their unique features in diverse human organs remains limited. To address this, we compiled a meta-analysis dataset of 114,275 CD45+ immune cells sourced from 14 organs in healthy donors. While the transcriptome of immune cells remains relatively consistent across organs, our analysis has unveiled organ-specific gene expression differences (GTPX3 in kidney, DNTT and ACVR2B in thymus). These alterations are linked to different transcriptional factor activities and pathways including metabolism. TNF-α signaling through the NFkB pathway was found in several organs and immune compartments. The presence of distinct expression profiles for NFkB family genes and their target genes, including cytokines, underscores their pivotal role in cell positioning. Taken together, immune cells serve a dual role: safeguarding the organs and dynamically adjusting to the intricacies of the host organ environment, thereby actively contributing to its functionality and overall homeostasis.
BackgroundPhenotypic data comparison is essential for disease association studies, patient stratification, and genotype-phenotype correlation analysis. To support these efforts, the Global Alliance for Genomics and Health (GA4GH) established Phenopackets v2 and Beacon v2 standards for storing, sharing, and discovering genomic and phenotypic data. These standards provide a consistent framework for organizing biological data, simplifying their transformation into computer-friendly formats. However, matching participants using GA4GH-based formats remains challenging, as current methods are not fully compatible, limiting their effectiveness.ResultsHere, we introduce Pheno-Ranker, an open-source software toolkit for individual-level comparison of phenotypic data. As input, it accepts JSON/YAML data exchange formats from Beacon v2 and Phenopackets v2 data models, as well as any data structure encoded in JSON, YAML, or CSV formats. Internally, the hierarchical data structure is flattened to one dimension and then transformed through one-hot encoding. This allows for efficient pairwise (all-to-all) comparisons within cohorts or for matching of a patient's profile in cohorts. Users have the flexibility to refine their comparisons by including or excluding terms, applying weights to variables, and obtaining statistical significance through Z-scores and p-values. The output consists of text files, which can be further analyzed using unsupervised learning techniques, such as clustering or multidimensional scaling (MDS), and with graph analytics. Pheno-Ranker's performance has been validated with simulated and synthetic data, showing its accuracy, robustness, and efficiency across various health data scenarios. A real data use case from the PRECISESADS study highlights its practical utility in clinical research.ConclusionsPheno-Ranker is a user-friendly, lightweight software for semantic similarity analysis of phenotypic data in Beacon v2 and Phenopackets v2 formats, extendable to other data types. It enables the comparison of a wide range of variables beyond HPO or OMIM terms while preserving full context. The software is designed as a command-line tool with additional utilities for CSV import, data simulation, summary statistics plotting, and QR code generation. For interactive analysis, it also includes a web-based user interface built with R Shiny. Links to the online documentation, including a Google Colab tutorial, and the tool's source code are available on the project home page: https://github.com/CNAG-Biomedical-Informatics/pheno-ranker.
The characterization of somatic genomic variation associated with the biology of tumors is fundamental for cancer research and personalized medicine, as it guides the reliability and impact of cancer studies and genomic-based decisions in clinical oncology. However, the quality and scope of tumor genome analysis across cancer research centers and hospitals are currently highly heterogeneous, limiting the consistency of tumor diagnoses across hospitals and the possibilities of data sharing and data integration across studies. With the aim of providing users with actionable and personalized recommendations for the overall enhancement and harmonization of somatic variant identification across research and clinical environments, we have developed ONCOLINER. Using specifically designed mosaic and tumorized genomes for the analysis of recall and precision across somatic SNVs, insertions or deletions (indels), and structural variants (SVs), we demonstrate that ONCOLINER is capable of improving and harmonizing genome analysis across three state-of-the-art variant discovery pipelines in genomic oncology.