
MOTIVATION: Phenotype concept recognition (CR) is a fundamental task in biomedical text mining. However, existing methods either require ontology-specific training, making them struggle to generalize across diverse text styles and evolving biomedical terminology, or depend on general-purpose large language models that lack necessary domain knowledge. RESULTS: To address these limitations, we propose AutoPCR, a prompt-based phenotype CR method designed to automatically generalize to new ontologies and unseen data without ontology-specific training. To further boost performance, we also introduce an optional self-supervised training strategy. Experiments show that AutoPCR achieves the best average and most robust performance across datasets. Further ablation and transfer studies demonstrate its inductive capability and generalizability to new ontologies. AVAILABILITY AND IMPLEMENTATION: Our code is available at https://github.com/yctao7/AutoPCR.
MOTIVATION:With the advancement of cryo-electron microscopy (cryo-EM) into the atomic resolution era, accurate Cα atom modeling has become essential for macromolecular structure determination. However, existing evaluation systems overly rely on full-atom metrics and lack a dedicated, comprehensive benchmark for assessing Cα prediction modules within automated modeling tools. RESULTS:To address this gap, we establish a rigorous benchmark to evaluate the Cα prediction performance of four prominent deep learning-based methods (ModelAngelo, DeepMainMast, EModelX, and CryoAtom) across multiple dimensions. We construct a diverse dataset covering a wide range of resolutions (1-8 Å), molecular weights, and noise levels. A novel evaluation framework is introduced, incorporating multi-threshold RMSD-based metrics (1-3 Å) alongside advanced point-cloud similarity measures (Chamfer Distance, Earth Mover's Distance) for quantitative and nuanced assessment. Our results reveal that method performance is highly dependent on the chosen evaluation criteria and intrinsic data characteristics. ModelAngelo excels under loose thresholds with high-quality data but shows sensitivity to resolution degradation; CryoAtom demonstrates notable computational efficiency, however, its completeness-oriented design leads to a certain loss of precision; EModelX demonstrates balanced generalization across varied conditions; DeepMainMast achieves high localization accuracy under stringent criteria but incurs a high computational cost. AVAILABILITY AND IMPLEMENTATION:This work provides a reproducible, Cα-centric evaluation framework to guide method development and advance automated cryo-EM structure determination. The source code for the benchmark and evaluation metrics is freely available at https://github.com/zhtianz/Benchmarking\_CA.
Motivation Single-cell sequencing technologies allow researchers to study cell-cell variation within a cell population. Variations between cells are driven by the underlying biological network, particularly gene regulatory networks (GRNs). GRNs rewire as cells evolve, and different cells can have different GRNs. However, while single-cell RNA-sequencing (scRNA-seq) and single-cell multi-omics data have been used to reconstruct GRNs, the output GRNs are rarely cell-specific, but rather, most existing methods infer population-level or cell-type-level GRNs.Results We propose CeSpGRN (Cell-Specific Gene Regulatory Network inference), a method that infers cell-specific GRNs from scRNA-seq, paired scRNA-seq and scATAC-seq, or spatial transcriptomic data. In particular, existing methods that use matching scRNA-seq and scATAC-seq data incorporate population-level region information in GRN inference, whereas CeSpGRN utilizes single-cell resolution region information. CeSpGRN infers cell-specific GRNs using a kernel-weighted Gaussian Copula Graphical Model, and incorporates multi-omic or spatial location information when constructing the objective function. We tested CeSpGRN on both simulated and real datasets, and the results show that CeSpGRN has a superior performance compared to baseline methods in reconstructing GRNs and detecting regulatory interactions that differ between cells. CeSpGRN uncovered regulatory interactions that rewire during biological processes on real datasets.Availability and implementation CeSpGRN is a Python package available at https://github.com/PeterZZQ/CeSpGRN.
Motivation Psoriasis is a chronic, immune-mediated disorder with an unmet need for effective treatments. To systematically prioritize therapeutic targets, we integrated proteome-wide Mendelian randomization (MR) with expression validation in blood/skin, genetic susceptibility analysis, differential gene expression (DGE) from bulk and single-cell RNA sequencing (scRNA-seq), colocalization, pathway enrichment, and protein-protein interaction analyses.Results Proteome-wide MR identified 29 candidate protein targets (Bonferroni-corrected), all replicated in independent datasets. Fifteen targets showed significant expression associations in blood or skin. Eleven proteins-UBLCP1, IL23A, ASF1A, RARRES2, ICAM1, PRSS53, ICAM5, GCA, IL2RA, DBI, and NFKB1-exhibited consistent directional effects with their genes. Genetic susceptibility analysis confirmed 20 target-specific polygenic scores for psoriasis and five for psoriatic arthritis. DGE analysis identified 13 targets in bulk and 13 in scRNA-seq-primarily in keratinocytes and immune cells-with IL2RA, COMP, and A2ML1 dysregulated across both. Colocalization analysis implicated shared causal variants for psoriasis in ASF1A, CD8A, CTF1, IL7R, MMP12, RARRES2, XCL2, DBI, IL23A, IL2RA, SGSH, and TIMD4. Enrichment analyses highlighted involvement in cytotoxicity, immune regulation, and JAK-STAT signaling. Eighteen targets interacted with approved anti-psoriasis drugs. Notably, drugs targeting IL2RA, IL7R, CTF1, ICAM1, MMP12, NFKB1, CD8A, DDX58, IL12A, SGSH, and FAP are approved or in trials for other diseases, suggesting repurposing potential. Our integrative multi-omics approach prioritized 29 high-confidence targets, including 13 novel candidates (RARRES2, ASF1A, CTF1, DBI, B3GNT2, CD8A, TIMD4, CRTAM, SGSH, XCL2, DAPK2, A2ML1, and FAP). Several high-priority targets-such as IL2RA, IL23, MMP12, RARRES2, IL7R, and ICAM1-were supported across analytical layers. These findings provide a robust foundation for psoriasis drug development.Availability and implementation The code used for the analyses in this manuscript has been archived in Zenodo at [DOI: 10.5281/zenodo.19692128].
Motivation: Data-enabled studies of microbial ecology and evolution depend on high-quality descriptions of microbial habitats, based on curated and consolidated vocabularies. Results: We introduce microntology v1.0, a pragmatic controlled vocabulary of 148 terms to describe microbial habitats and lifestyles, and provide manually curated microntology annotations for >300k metagenomic samples from public repositories. Availability: microntology controlled vocabulary terms and term hierarchies (doi: 10.5281/zenodo.19730167), and curated annotations for 305 626 metagenomic samples (doi: 10.5281/zenodo.18164252) are available via Zenodo and spire.embl. de/downloads. Underlying code is available via github.com/grp-schmidt/microntology and Zenodo (doi: 10.5281/zenodo. 20323497). User feedback, suggestions and bug reports are welcome at github.com/grp-schmidt/microntology/issues.
Motivation Accurately identifying compound-protein interactions (CPIs) is critical for accelerating drug discovery. Recent deep learning methods have achieved impressive results, yet they primarily focus on local structures and neighborhood information, often overlooking high-order interaction patterns shared among similar molecules.Results In this paper, we propose HKD-CPI, a high-order knowledge-enhanced inductive framework designed to improve generalization to unseen compound-protein pairs. Specifically, HKD-CPI introduces a molecular graph tokenization mechanism that aligns compound molecular graph features with token embeddings from sequence-pretrained large language models (LLMs), effectively infusing sequence-derived semantics into structural representations. To capture shared interaction patterns among functionally similar biomolecules, we construct a hypergraph-based representation to model high-order relationships between feature-similar compound/protein groups and their binding partners. Furthermore, a knowledge distillation strategy is further adopted to transfer high-order interaction knowledge from the hypergraph to a lightweight student model, enabling efficient and robust CPI prediction. Extensive experiments demonstrate that HKD-CPI outperforms existing state-of-the-art methods in inductive CPI prediction tasks. In particular, it achieves an average improvement of 4.94% in AUROC and 3.64% in AUPRC over the best-performing baseline across five benchmark datasets.Availability and implementation Our code and data are available at https://github.com/Hezy618/HKD-CPI.
SUMMARY:We developed slideimp, an R package that extends and optimizes K-nearest neighbor (K-NN) and Principal Component Analysis (PCA) imputation with grouped and sliding-window modes for accurate and efficient imputation of microarray and whole-genome DNA methylation (DNAm) data, respectively. Under a realistic scenario, slideimp achieved ≈12-28× faster runtime and ≈3-6× peak memory usage reduction for DNAm microarray imputation (GSE286313, EPICv2, N = 72) and achieved high imputation accuracy in a whole-genome DNAm dataset (N = 41). AVAILABILITY AND IMPLEMENTATION:The code used in this study is available at https://github.com/hhp94/slideimp_paper. The R package slideimp is available on CRAN (DOI: 10.32614/CRAN.package.slideimp). Version 1.0.0 of slideimp, which was used in this study, is archived on Zenodo (DOI: 10.5281/zenodo.20029382).
Motivation Chromatin regulation is crucial for modulating gene expression and cellular function by altering DNA accessibility. Defining and understanding chromatin regulation across diverse biological conditions, including health and disease, requires quantification of both the presence and enrichment level of diverse DNA-binding factors and chromatin modifications across defined genomic regions. Existing approaches mainly rely on peak-based or genome-wide models, which identify high-signal regions but do not annotate chromatin status at predefined functional genomic regions, such as promoters or enhancers. This lack of region-based annotation limits downstream comparative and integrative analyses across multiple factors and datasets, prompting us to create ChromCall.Results ChromCall is an R package for region-based chromatin enrichment analysis that provides a robust and extensible foundation for transparent and reproducible epigenomic profiling at predefined genomic regions. We applied ChromCall to ChIP-seq data from glioblastoma (GBM) brain tumours and found that the promoters of genes implicated in treatment resistance are significantly more likely to exhibit a combination of histone marks associated with phenotypic plasticity. This highlights a potential novel mechanism of therapeutic escape in these deadly tumours.Availability and Implementation The R package is available on https://github.com/GliomaGenomics/ChromCall and the version used in this paper is archived at https://doi.org/10.5281/zenodo.19580967
MOTIVATION:Approximate string matching (ASM) is the problem of finding all occurrences of a pattern in a text while allowing up to k errors. Many modern methods use seed-chain-extend, which is fast in practice, but does not guarantee finding all matches with ≤k errors. However, applications such as CRISPR off-target detection require exhaustive results. RESULTS:We introduce Sassy, a library and tool for ASM of short patterns in long texts. Sassy splits the text into four parts that are searched in parallel, and uses bitvectors in the text direction rather than the pattern direction. This has complexity O(k⌈n/W⌉) when searching a random text of length n, where W=256 is the SIMD width, and provides significant speedups for small k. Separately, we allow matches of the pattern to extend beyond the text for an overhang cost of, e.g. α=0.5 per character, to find matches near contig or read ends.Sassy is 4× to 15× faster than Edlib for patterns ≤1000 bp, and can search text with a throughput near 2 Gbp/s. Likewise, Sassy is over 100× faster than parasail. We apply Sassy to CRISPR off-target detection by searching 61 guide sequences in a human genome. Sassy is 100× faster than SWOffinder and only slightly slower (for k≤3) than CHOPOFF, for which building its index takes 20 min. Sassy also scales well to larger k, unlike CHOPOFF whose index took over 10 h to build for k=5. AVAILABILITY AND IMPLEMENTATION:Sassy is available as library and binary at https://github.com/RagnarGrootKoerkamp/sassy, and archived at swh:1:dir:e884758dce5777a441bc2799dc8824e563c5f97b.
MOTIVATION:GEDI is a generative framework for multi-sample, multi-condition single-cell analysis that performs batch correction, latent representation learning, and clustering-free differential expression within a unified model. However, the original implementation suffered from prohibitive memory use and runtime, preventing its application to modern atlas-scale datasets. RESULTS:We present GEDI 2.0, a complete high-performance reimplementation featuring a standalone C++ computational core with pre-allocated workspaces, strict sparse-matrix preservation, optimized BLAS routines, and multi-threaded block-coordinate descent. Across extensive benchmarks spanning up to 500 000 cells and 10 000 features, GEDI 2.0 achieves 40%-63.6% mean reduction in peak memory, 2.98× mean single-threaded speedups, and up to 11.5× acceleration with parallel execution, while maintaining full numerical equivalence to the original method. These improvements enable GEDI 2.0 to analyze million-cell datasets, a scale not achievable with the legacy implementation. GEDI 2.0 provides R and Python interfaces and seamless interoperability with common single-cell workflows. AVAILABILITY AND IMPLEMENTATION:Source code, documentation, reproducible codebase, and tutorials are available at https://github.com/csglab/gedi2.
MOTIVATION:Trimmomatic is a widely adopted tool for preprocessing high-throughput sequencing data, particularly from Illumina platforms. Since its original publication in 2014, the volume and complexity of sequencing data have increased dramatically, necessitating continuous tool evolution. RESULTS:We present the substantial updates to Trimmomatic over the past decade. Key enhancements include a robust multithreading model for high-performance parallel processing, parallel GZIP/BZIP2 compression, and a suite of new trimming and filtering steps to provide users with more flexible quality control. Usability has been significantly improved through automatic PHRED encoding detection and simplified file handling. The codebase has also been modernized including Maven support, and continuous integration to ensure long-term sustainability and community contributions. These updates solidify Trimmomatic's role as an efficient, flexible, and essential tool in modern bioinformatics pipelines. AVAILABILITY:Trimmomatic remains open-source under the GPL V3 license, with the latest version available at https://github.com/usadellab/Trimmomatic and also on our website https://www.plabipd.de/trimmomatic_main.html (DOI: https://doi.org/10.5281/zenodo.18678155).
Motivation Named entity recognition (NER) is a fundamental component of structured knowledge extraction, yet its effectiveness in emerging domains remains by the scarcity of high-quality, domain-specific annotated corpora. Although data augmentation and distant supervision have been explored to alleviate this issue, existing methods often introduce limited entity diversity, noisy labels, or disrupt contextual integrity, thereby limiting their generalization ability in low-resource settings.Results In this study, we propose DA-BioNER, a context-preserving data expansion framework for biomedical NER. DA-BioNER combines multiple base NER models trained on few-shot data to provide coarse annotations, followed by refinement using a large language model (LLM) guided by global biomedical knowledge. Unlike generation-based augmentation methods that synthesize new sentences, DA-BioNER performs annotation refinement within existing sentences, preserving both syntactic structure and semantic context. By constraining the role of LLM to refinement rather than open-ended generation, the framework effectively reduces hallucination while improving label precision and consistency. We evaluate DA-BioNER on three benchmark datasets (NCBI-Disease, BC5CDR, and BioRED), under low-resource conditions. In 40-shot settings, DA-BioNER achieves F1-scores of 0.750, 0.795, and 0.799, respectively, outperforming state-of-the-art methods, including LSMS, DAGA, and MELM, by up to 0.32. Under more extreme few-shot settings, DA-BioNER further improves F1-scores by up to 0.08, while generating an average of 1,391 additional unique entities, substantially enriching training diversity. These results demonstrate that DA-BioNER provides a scalable and adaptable solution for robust biomedical NER, particularly in domain adaptation and low-resource scenarios.Availability DA-BioNER is publicly available at https://github.com/DMnBI/DA-BioNER.
Motivation Serial section electron microscopy (ssEM) is essential for studying biological cell structures at nanometer resolution. However, supporting film folding (SFF) degradation frequently occurs during sample preparation, causing structural distortions and information loss that severely impair downstream analyses such as 3D reconstruction and neuron segmentation.Results We propose RegInpaint, a novel recovery framework that jointly addresses deformation correction and missing-information restoration caused by SFF degradation. RegInpaint formulates SFF recovery as a joint problem of 3D elastic registration and image inpainting, providing a generalizable solution for ssEM restoration. Experiments on four EM datasets show that RegInpaint consistently outperforms existing methods in image restoration quality and significantly improves neuron segmentation accuracy.Availability and implementation Source code is freely available at https://github.com/zhangzhenbang2021/RegInpaint.git.
Motivation Cryo-electron microscopy (Cryo-EM) single particle analysis (SPA) is a key technique for revealing the structure of biomacromolecules by three-dimensional reconstruction. Achieving high-resolution reconstruction relies on the acquisition of a large number of authentic particles; however, manual particle picking is inefficient and inadequate for the demands of reconstruction, making automated particle picking a major research focus. Although the foundational segmentation model Segment Anything Model (SAM) has recently advanced automated particle picking, its segmentation advantages have not been fully realized in cryo-EM applications. Moreover, cryo-EM images often have significant noise. Conventional denoising decreases noise but frequently overlooks high-level semantic information, leading to oversmoothed particle regions and reduced particle distinguishability.Results To address these challenges, we propose CryoPromptSeg, which employs prompt-guided SAM for particle picking while integrating a semantically enhanced image denoiser. Specifically, by performing domain adaptation fine-tuning of SAM and incorporating prompts generated by the proposed automatic prompt generator, it achieves precise segmentation of cryo-EM particles. In addition, it employs a parallel multi-task framework to jointly train the denoiser and the prompt generator, incorporating particle semantic information from the prompt generator into the denoiser to suppress noise while preserving highly distinguishable particle structures. To lower the barrier to practical application, we developed a user-friendly online prediction platform for particle picking. Experimental results demonstrate that CryoPromptSeg outperforms existing mainstream methods in both particle picking accuracy and image denoising quality, thus providing a novel solution for the automation of particle picking.Availability The code and platform are available at: https://github.com/347251369/CryoPromptSeg.
Motivation Electrophysiological recordings are essential in experimental and computational neuroscience, providing insights into neuronal excitability and network behaviour. Extracting features such as action potential thresholds, widths, and firing patterns is conceptually straightforward, but in practice it is complicated by heterogeneous datasets and software environments, which hinder reproducibility and interoperability. A standardized, efficient, and portable framework is needed to ensure consistent analysis across platforms and alignment with community data standards.Results We present the Electrophysiology Feature Extraction Library (eFEL), a cross-platform, open-source library that implements standardized definitions for over 90 electrophysiological features. eFEL combines a high-performance C++ core with a Python interface, supporting customizable feature dependencies, caching, and parallelization. It integrates with community standards such as Neurodata Without Borders and works seamlessly with common electrophysiology formats and simulation environments. Since its initial release in 2015, eFEL has been used in published studies spanning single-cell analysis, model optimization, multimodal fitting, and circuit simulations. eFEL provides a FAIR-compliant, versatile resource for reproducible electrophysiological data analysis.Availability and implementation The eFEL library is publicly available at https://github.com/openbraininstitute/eFEL and the associated study data and scripts have been deposited in Zenodo at https://zenodo.org/records/17241835.
SUMMARY:Although RNA-sequencing has replaced microarrays for gene expression profiling over the past 15 years, its full potential for splicing analysis in clinical settings remains underexploited. Most available tools are tailored for large cohorts or known isoforms, limiting their applicability in routine diagnostics where non-recurring events must be identified in low-dimension datasets. We present SAMI (Splicing Analysis with Molecular Indexes), a fully-integrated UMI-aware pipeline designed to detect splicing events diverging from transcript annotations. Building upon the well-proven STAR aligner, SAMI introduces original post-processing of gaps and potential intron retention to maximize accuracy, along with clear graphical representations and tunable filtering stringency. The ability of SAMI and concurrent software to detect intragenic splicing aberrations and gene fusions was assessed, both on real data from a commercial control sample and simulated data generated with ASimulatoR. AVAILABILITY AND IMPLEMENTATION:Nextflow pipeline and Singularity container recipe freely available under GPL 3 license at https://github.com/HCL-HUBL/SAMI.
Motivation The accurate and sensitive identification of de novo variants, which are unique to an individual and not found in the parents' germlines, is critical for understanding the genetic basis of rare diseases, developmental disorders, and evolutionary processes. Existing de novo variant detection pipelines often lack the flexibility to handle multiple variant types, struggle with speed and reproducibility across computational environments, demand extensive manual configuration, or require bioinformatics expertise for downstream curation and analysis, limiting their scalability and usability for large genomic studies. Accordingly, there is a pressing need to better address these challenges.Results We introduce TriosCompass, an open-source Snakemake workflow that addresses these challenges by providing a modular, accelerated, and environmentally-configurable end-to-end solution for comprehensive de novo variant discovery. It integrates state-of-the-art tools into a reproducible framework, empowering researchers to discover novel genetic insights with greater efficiency and reliability.Availability TriosCompass is implemented as a Snakemake workflow and is freely available at https://github.com/NCI-CGR/TriosCompass_v2 or on Zenodo (10.5281/zenodo.17981062).Supplementary information Supplementary data is available on GitHub at https://github.com/NCI-CGR/TriosCompass_v2/tree/manuscript/report_dashboards. Supplementary methods on DeepTrio benchmark runs can be viewed at: https://github.com/NCI-CGR/TriosCompass_v2/blob/manuscript/TriosCompass_Supp_Methods_deeptrio_benchmark.md
Motivation Large-scale omics resources, including The Cancer Genome Atlas, Genomics of Drug Sensitivity in Cancer, and the Cancer Dependency Map, have become essential for cancer research. However, these datasets are distributed across different platforms, formats and analysis frameworks, which limits their practical use by researchers without extensive computational expertise.Results We developed CancerOmicsStudio (CoS), a web server for integrative and interpretable analysis of multi-omics cancer data across 33 cancer types. CoS provides five major modules: CosAI, Traditional Analysis, Drug Sensitivity, CRISPR Dependency and Single-Cell Tumor Microenvironment. The Traditional Analysis module supports expression comparison, diagnostic evaluation, survival analysis, enrichment analysis and gene correlation. The Drug Sensitivity and CRISPR Dependency modules enable systematic evaluation of gene-drug response associations and gene essentiality in cancer cell lines. The Single-Cell Tumor Microenvironment module supports tumor microenvironment analysis at single-cell resolution. In total, approximately 1.23 million results have been precomputed to enable rapid retrieval. CosAI further allows users to submit natural-language queries and obtain results through a Real-time Analysis as Retrieval framework, with responses summarized by a lightweight language model.Availability and implementation CancerOmicsStudio is freely available at Zenodo (doi: 10.5281/zenodo.18744990) and https://cos.wanglab.bio.
MOTIVATION:Network analysis has become a central strategy for dissecting complex biological and environmental systems, particularly as modern omics technologies generate increasingly large and heterogeneous datasets. However, current tools often lack the scalability, flexibility, and native multi-omics support required for high-dimensional data analysis. We developed MetaNet, a high-performance R package that unifies network construction, visualization, and analysis across diverse omics layers. RESULTS:MetaNet enables fast and scalable correlation-based network construction for datasets with more than 10 000 features, providing over 40 layout algorithms, rich annotation utilities, and visualization options compatible with both static and interactive platforms. It further offers comprehensive topological and stability metrics for in-depth network characterization. Benchmarking shows that MetaNet delivers up to a 100-fold improvement in computation time and a 50-fold reduction in memory usage compared to existing R packages. We demonstrate its utility through two representative applications: (1) longitudinal microbial co-occurrence networks revealing airborne microbiome dynamics, and (2) an integrative exposome-transcriptome network of over 40 000 features, uncovering distinct regulatory impacts of biological and chemical exposures. By offering a robust, reproducible, and biologically informed framework, MetaNet advances multi-omics network analysis across biological, ecological, and environmental domains. AVAILABILITY:MetaNet package is freely available at https://github.com/Asa12138/MetaNet.
Motivation Cell-cell interactions (CCIs) are fundamental to multicellular organisms and play crucial roles in diverse biological processes and disease mechanisms. Understanding CCIs is vital for deciphering disease pathogenesis and developing therapeutic strategies. Although numerous computational methods have been developed to infer CCIs from complex biological data, most existing approaches rely primarily on single-gene expression levels and ligand-receptor databases, often failing to capture the nuanced network-wide changes characteristic of disease states.Result We propose MetaCCI, a novel computational strategy that integrates meta-information into CCI inference by extending the traditional gene expression-based analysis to a gene regulatory network framework. MetaCCI meticulously combines established ligand-receptor pairs with quantitative insights into gene behavior within complex gene networks, enabling the precise extraction of relevant targets for CCI inference. Subsequently, CCI inference was performed using an eigen cell co-expression network, providing a more holistic view of cell-cell communication. Monte Carlo simulations demonstrated that MetaCCI consistently outperforms existing methods in CCI inference. We applied MetaCCI to characterize cell-cell communication in Myelodysplastic Syndromes (MDS). Our results identified distinct interaction patterns in MDS compared with normal cell populations, specifically highlighting the loss of CCIs between "Dendritic cells and Hematopoietic precursor cells" and between "Dendritic cells and Hematopoietic multipotent progenitor cells" as characteristic features of MDS. Furthermore, FABP5, CD63, and HMGB1 were identified as MDS-specific markers. These findings suggest that diminished CCIs involving dendritic cells, hematopoietic precursor cells, and multipotent progenitor cells are pivotal to MDS pathogenesis.Availability and implementation The MetaCCI software is freely available at https://github.com/HeewonGitHub/MetaCCI. An archived version of the software and example datasets used in this study is available at Zenodo: https://doi.org/10.5281/zenodo.20101527.