In alignment with the European Health Data Space (EHDS), Spain's IMPaCT - Precision Medicine Infrastructure associated with Science and Technology - aims to establish a trusted research environment (TRE) for secure and FAIR data sharing and analysis. It is structured into three pillars: Predictive Medicine (IMPaCT-Cohort), Genomic Medicine (IMPaCT-Genomics), and Data Science (IMPaCT-Data). IMPaCT-Data leads the development of the IMPaCT Digital Platform (IDP), integrating clinical, genomic, and imaging data to support the national IMPaCT-Cohort and Personalised Medicine Projects (IMPaCT-PMPs). Its Reference Implementation defines the architecture across a federated model.
Therapeutic synergy emerges from interactions between molecular drug action, intracellular signaling, and tissue-level transport dynamics. We developed a multiscale model integrating these scales to predict schedule-dependent drug combination effect in the AGS cell line. Calibrated solely on single-drug growth curves, the model accurately predicted population-level outcomes of drug combinations without combination-specific training. This demonstrates the model’s capacity to suggest mechanistic multiscale insights into the logic of drug combinations within the AGS cell line, establishing a computational platform for the systematic in silico exploration of virtual multiscale experiments on drug diffusion and dosing schedules. Cross-scale analysis revealed that combination therapy efficacy flows across scales: population-level pharmacokinetics dictate the sequence of molecular target engagement within individual cells, determining collective cell-fate decisions. Our simulations predict that inhibiting the PI3K/AKT axis before MEK is more effective than the reverse order, disabling a pro-survival rebound and locking cells into an apoptotic state. Ultimately, by capturing phenomena that single-scale approaches cannot, this framework generates translationally relevant hypotheses, providing a versatile platform for optimizing drug scheduling, formulation, and combination strategies where matched molecular Boolean models and phenotypic data are available.
The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/) in response to the need for rapid data sharing and analysis during the 2020-2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. It considers useful lessons and approaches for future pandemic preparedness which enable researchers to easily identify and obtain the key data from resources on a national and international level.
Deep learning methods, including deep representation learning (DRL) approaches such as variational autoencoders (VAEs), have been widely applied to cancer omics data to address the high dimensionality of these datasets. Despite remarkable advances, cancer is a complex and dynamic disease, making it challenging to study, and the temporal resolution of cancer progression captured by omics-based studies remains limited. In this systematic literature review, we explore the use of DRL, particularly the VAE, in cancer omics studies for modeling time-related processes, such as tumor progression and evolutionary dynamics. Our work reveals that these methods most commonly support subtyping, diagnosis, and prognosis in this context, but rarely emphasize temporal information. We observed that the scarcity of longitudinal omics data currently limits deeper temporal analyses that could enhance these applications. We propose that applying the VAE as a generative model to study cancer in time, particularly focusing on cancer staging, could lead to meaningful advancements in our understanding of the disease.
The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/)inresponsetothe need for rapid data sharing and analysis during the 2020–2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. In this study, we observe that pandemic response greatly benefits from an established infrastructure that can be quickly mobilised, developed and extended. Collaborations and preparation built on solid foundations over several years, supported by investment in the form of national and international research grants, is key for sustainability, continuation and readiness to deploy such efforts.
Predicting cancer stage progression from omics data, and deriving molecular insight into the mechanisms driving it, remains a major challenge, owing in part to the lack of adequate longitudinal data and the interpretability limitations of current forecasting models. Large cancer datasets such as TCGA capture patient profiles cross-sectionally rather than longitudinally, complicating timely treatment decisions as tumors become more invasive. Deep neural networks typically used for forecasting, such as LSTMs, compound this problem by remaining largely opaque and offering clinicians no straightforward way to audit their predictions. Clear cell renal cell carcinoma (ccRCC) illustrates the clinical stakes of both challenges. Five-year survival falls from over 94% at stage I to 28% at stage IV, yet early-stage tumors are often managed under active surveillance, a strategy constrained by sparse molecular evidence of progression risk. Detecting progression in time, meanwhile, demands forecasts clinicians can interpret and trust, not black-box predictions. We address both challenges by combining generative and symbolic AI: a Variational Autoencoder trained on bulk RNA-Seq profiles of 530 TCGA ccRCC patients generates synthetic pseudo-time trajectories that overcome the absence of longitudinal data, while a symbolic rule-induction framework (ASAL) learns finite-state automata from these trajectories, encoding stage transition as human-readable Boolean conditions over gene expression, which a complex event forecasting system (Wayeb) converts into probabilistic forecasts of stage advancement. An independent XGBoost classifier trained on real patients (F1 score = 0.71-0.81) shows a gradual early-to-late probability shift along the synthetic trajectories, absent in non-progressing control trajectories. Pathway enrichment of those trajectories reveals stage-dependent changes in established kidney cancer-related processes, including the TCA cycle and DNA repair. Finally, our symbolic forecaster nearly matches an LSTM baseline (macro F1 = 0.928 vs. 0.964), while additionally offering an inspectable rule set and a probability distribution over transition timing rather than a single opaque score. This work shows that generative and symbolic AI, paired together, can turn cross-sectional cohorts into a transparent, forecast-oriented framework for modeling disease progression, demonstrated here in ccRCC.
OBJECTIVES:Report the development of the patient-centered myAURA application and suite of methods designed to aid epilepsy patients, caregivers, and clinicians in making decisions about self-management and care. MATERIALS AND METHODS:myAURA rests on an unprecedented collection of epilepsy-relevant heterogeneous data resources, such as biomedical databases, social media, and electronic health records (EHRs). We use a patient-centered biomedical dictionary to link the collected data in a multilayer knowledge graph (KG) computed with a generalizable, open-source methodology. RESULTS:Our approach is based on a novel network sparsification method that uses the metric backbone of weighted graphs to discover important edges for inference, recommendation, and visualization. We demonstrate by studying drug-drug interaction from EHRs, extracting epilepsy-focused digital cohorts from social media, and generating a multilayer KG visualization. We also present our patient-centered design and pilot-testing of myAURA, including its user interface. DISCUSSION:The ability to search and explore myAURA's heterogeneous data sources in a single, sparsified, multilayer KG is highly useful for a range of epilepsy studies and stakeholder support. CONCLUSION:Our stakeholder-driven, scalable approach to integrating traditional and nontraditional data sources enables both clinical discovery and data-powered patient self-management in epilepsy and can be generalized to other chronic conditions.
Direct and inverse comorbidities between neuropsychiatric disorders and cancer are increasingly recognised as important features of the nervous system-cancer relationship, yet the inherited genetic architecture underlying these patterns remains poorly understood. Here, we analysed pairwise genetic correlations across 35 diseases represented by 115 GWAS datasets, including 9 psychiatric disorders, 10 neurological diseases and 16 cancers, using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL), complemented by meta-analysis, subtype-resolved analyses, local covariance mapping and multi-omic benchmarking. Genetic correlations were predominantly positive and strongest within disease categories, whereas cancer-neurological pairs showed the weakest overall genetic affinity. Meta-analysis and subtype resolution uncovered associations obscured in aggregate analyses, including opposing correlations between familial and late-onset Alzheimer's disease and lung cancer, revealing subtype-dependent neuro-oncological biology. Local analyses identified recurrent genomic loci where direct comorbidities are consistent with shared inflammatory, interferon, survival and tissue-remodelling programs, whereas inverse comorbidities suggest competing demands on apoptotic regulation, immune tone and stress-response calibration between neuronal and tumour-cell states. Together, these findings provide a genome-scale genetic framework for neuropsychiatric-cancer comorbidity and identify shared inherited biological programs as candidates for mechanistic investigation and therapeutic translation.
The Spanish National Bioinformatics Institute (INB), founded in 2003 as a distributed network, is the ELIXIR Node in Spain and has two objectives: 1) deepen its involvement and leadership within ELIXIR and broaden the resources provided as part of ELIXIR infrastructure to the Life Sciences community; and 2) increase its impact within the Spanish National Health System . INB/ELIXIR-ES continues to strengthen its technological capabilities in federated data infrastructures, interoperability, and FAIR data management within ELIXIR. The Node is actively involved in several ELIXIR-driven projects and commissioned services (CoS) within the 2024-28 Work Programme. From the ELIXIR perspective, the Service Delivery Plan (SDP) maintains 40 resources offered by 24 groups belonging to 12 institutions. Regarding national activities, the INB/ELIXIR-ES leads the Translational Bioinformatics Network (TransBioNet) and serves as proxy between IMPaCT-Data activities, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine, and European efforts. These activities align with major European data projects such as the Genomic Data Infrastructure (GDI), EUCAIM, and the Federated European Genome-phenome Archive (FEGA), while implementing Global Alliance for Genomics and Health (GA4GH) standards in its technological developments. The Node strongly engages within the 2024-28 Work Programme: TechnologyTier: co-leadership of Data, Tools and Training Platforms, with contributions across all Platforms. During this period, the co-led ELIXIR Beacon Network Infrastructure Service secured funding for this service, strengthening federated data discovery capabilities. ScienceTier : co-leadership of CMR and HDTR Science priority areas; co-leadership of Rare Diseases, FHD, Cancer Data and Biodiversity Communities; and Pathogens Data, and RNA Data Focus Group; leading and participating in several CoS in the HDTR and CMR areas. PeopleTier : co-leadership of the ELEAD2.0 leadership programme, a CoS built on the experiences of Bioinfo4Women, and active role in the PeoplePulse CoS. Active role in the NodeTier NSCS, together with 4 Platforms, 13 Communities, and 8 Focus Groups. Regarding the INB/ELIXIR-ES portfolio, EGA, an ELIXIR CDR, is co-developed and maintained by CRG and EMBL-EBI with BSC’s infrastructure support, with the current focus on its extension through Federated EGA. Canada joined the Federated EGA, marking the first major expansion beyond Europe and reinforcing its global dimension. Additionally, four resources are recognised as ELIXIR RIRs: 3DBIONOTES-API, FAIRtracks, FAIRCookbook and OpenEBench. Various ELIXIR Communities have adopted OpenEBench as their community-driven benchmarking platform. The INB/ELIXIR-ES continued its training activities and organised key meetings within ELIXIR. It gathered its national community in the XV Symposium on Bioinformatics (JBI2025) jointly organised with ELIXIR-PT and INSTRUCT-ES. Other ELIXIR events organised were the ELIXIR 3DBioinfo Community Annual General Meeting with the 3D-SIG Community, and the Biodiversity and Microbiome Community meetings. https://inb-elixir.es https://inb-elixir.es/resources
Abstract Complex diseases are influenced by both genetic and environmental factors. Immune cells are key mediating interactions with the environment, but the impact of genetic variation on the immune system and how it influences complex diseases is not fully understood. Moreover, most genome-wide analyses (GWAS) variants associated with complex diseases are non-coding and difficult to interpret. Here, we investigated the association of non-coding variants with immune cell enhancers. As part of BLUEPRINT and the International Human Epigenome (IHEC) consortia, we generated and analysed a comprehensive set of epigenomes for human primary immune cells, including 107 epigenomes derived from 749 ChIP-Seq experiments across 24 cell types. We identified multicell enhancer activity patterns across the genome and examined their links with non-coding variants from 518 GWAS traits. This analysis revealed 117 significant associations, including novel links between cardiovascular disease variants and macrophage-specific enhancers that regulate genes involved in lipid metabolism and immunity, such as the gene encoding for the nuclear receptor LXR-alpha ( NR1H3) and many of its known target genes. Together, these data will help to better understand the influence of genetic variability in immune function and related diseases.
ABSTRACT IMPaCT-Data, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine coordinated by the Barcelona Supercomputing Center (BSC), has assembled and deployed a national-scale infrastructure for clinical, genomics, and imaging workloads across major Spanish biomedical research institutions. The main technical challenge is enabling large-scale, compute-intensive workflows while complying with strict data-governance constraints, heterogeneous local infrastructures, and non-uniform compute and storage capabilities. The first layer of the platform is based on a distributed execution model in which workflows are centrally dispatched while computation is performed locally at participating institutions. A central Galaxy server instance hosted at BSC provides workflow management, provenance tracking, user access, and operational monitoring. Authentication and authorization are implemented using OpenID Connect (OIDC), with institutional identity providers federated through a central Keycloak service, allowing the preservation of local identity management policies. Galaxy Pulsar handles job dispatching and, based on asynchronous messaging (RabbitMQ/AMQPS), distributes workloads to remote federated nodes and delegates processing to site-local backends (e.g. container engines, HPC clusters). Analysis tools and reference datasets are distributed and synchronized through CVMFS, reducing the operational burden associated with maintaining synchronized analysis environments across a federated infrastructure. For highly sensitive datasets, a second execution network is being deployed, enabled and orchestrated using the Federated Execution Manager (FEM), and exposed to the researcher through a tailored virtual research environment powered by openVRE (future safeVRE). Under this architecture, participating institutions will retain full control not only over data access but also over the execution environment and approved analysis tools, making it particularly suitable for secure processing environments (SPEs). In this context, FEM allows institutions to enforce local governance policies, validate containers and workflows before execution, and control how intermediate results are generated and shared across sites. This design also supports regulated biomedical scenarios in which datasets remain in read-only mode and computation must be executed under strict auditing and traceability requirements. In addition, FEM natively supports more complex federation patterns, including federated analysis (calculator-aggregator schemes) and federated learning workflows (client-server schemes). The resulting platform provides a practical technical blueprint for biomedical data processing aligned with ELIXIR principles and other major efforts and initiatives, such as Genomics Alliance for Genomics and Health (GA4GH), EUCAIM, and GDI, among others: distributed execution close to data, standards-aligned interoperability, and reproducible workflows across a clinically governed federated network.
The rapid evolution of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has compromised the efficacy of many authorized monoclonal antibody products. This highlights the need for alternative strategies, especially for vulnerable populations such as immunocompromised individuals. Here, we optimized angiotensin-converting enzyme 2 (ACE2)-Fc fusion proteins by combining three engineering steps: in silico mutagenesis of the S protein binding interface to increase affinity, insertion of a flexible linker to improve protein stability and S protein accessibility, and generation of a tetrameric molecule to maximize avidity. Neutralizing activity was tested against a large panel of pre-Omicron and Omicron pseudoviruses and authentic viruses, including JN.1 and KP.2 variants. Optimized ACE2-Fc molecules demonstrated potent neutralizing activity, in the picomolar range, against all SARS-CoV-2 variants. Our molecules displayed similar potency but better resilience when compared to the monoclonal antibody Sipavibart. These findings support ACE2-Fc proteins as robust candidates for next-generation interventions against infection by an evolving SARS-CoV-2.
T cell receptor (TCR) recognition of peptide-MHC complexes (pMHCs) is central to adaptive immunity. Structural insights into TCR-pMHC interactions are critical for understanding antigen specificity and T-cell function. However, progress remains limited by the scarcity of experimentally resolved structures (275 TCR-pMHC class I structures in the PDB, Jan 2026). Although protein structure modelling tools have advanced rapidly, accurate structural modelling of TCRs remains challenging due to CDR loop hypervariability and conformational flexibility. In addition, there is a lack of reliable quality assessment strategies that do not rely on comparisons with experimental references. To address this, we benchmarked four general-purpose (AlphaFold2.3-Multimer, AlphaFold3, Boltz-2, Chai-1) and three TCR-specific (TCRmodel2, tFold-TCR, TCRdock) protein modelling algorithms by recalculating all experimentally determined TCR-pMHC class I complexes in the PDB. AlphaFold3 demonstrated superior performance across metrics (mean TCR-iRMSD = 3.59 Å and DockQ = 0.54), whereas the other algorithms displayed lower accuracy. Built on AlphaFold3 structures, we present a scalable and interpretable ML framework for the quality assessment of TCR-pMHC structural models without matched experimental references. We trained a random forest classifier integrating multiple confidence metrics (pLDDT, ipTM, ipSAE, iPAE, iPDE, pDockQv1-2) derived from 1325 modelled structures of 265 experimentally determined PDB TCR-pMHC class I complexes. The classifier reliably stratifies structural models into low-, acceptable-, medium-, and high-quality tiers defined by comparisons to their experimental reference structures, outperforming single metrics. These quality-tier predictions further enable the prioritization of high-confidence TCR-pMHC interactions. This was demonstrated across two held-out datasets comprising a total of 4,090 AlphaFold3-modelled TCR-pMHC complexes (20,450 models, 5 models per complex): a re-evaluated set of TCR-pMHC class I complexes from VDJdb (n = 606) and the reference IMMREP23 dataset (n = 3,484). We profiled the first dataset to reduce false-positive TCR-pMHC interactions erroneously annotated in VDJdb, and the second to enrich for biologically validated TCR-pMHC interactions amongst higher-quality structural models over their synthetic negative counterparts. Altogether, our structural quality-tier framework provides a scalable and interpretable approach that complements structural modelling and functional analyses of TCR-pMHC class I complexes, with direct translational applications in T-cell immunology and TCR-based immunotherapies.
Abstract Genome-scale metabolic models can predict how individual cells allocate resources and respond to their environment, yet few frameworks link single-cell metabolism to the spatial organisation of multicellular systems. Here we introduce PhysiCelldFBA, an extension of the PhysiCell agent-based framework that couples genome-scale dynamic flux balance analysis to off-lattice multicellular simulations. Each simulated cell carries its own metabolic model, allowing local environmental conditions to shape metabolism while metabolic activity feeds back on the surrounding environment, cellular behaviour, and spatial organisation. We first validate this coupling by showing that glucose consumption, CO 2 production, and biomass accumulation remain mass-balanced in a closed E. coli system, with simulated biomass agreeing with analytical predictions to within 1%. We then demonstrate how metabolic phenotypes emerge from this coupling across microbial and mammalian systems. Spatial nutrient gradients generate metabolic stratification and acetate cross-feeding in growing E. coli colonies; diffusion-limited metabolism produces proliferative, hypoxic, and necrotic zones across a broad panel of metabolites in a tumour-like tissue; distinct, organism-specific metabolic networks give rise to syntrophic cross-feeding and spatial niche formation in a two-species consortium; and metabolic state couples energy availability to transitions between cellular motility and growth. Across these examples, metabolic stratification, cross-feeding, and phenotypic adaptation emerge from local metabolic optimisation and environmental feedback rather than being explicitly prescribed. PhysiCelldFBA therefore provides a general framework for simulating genome-scale metabolism at single-cell resolution and linking intracellular metabolic state to cellular behaviour and emergent organisation across scales.
Agent-based cellular models simulate tissue evolution by capturing the behavior of individual cells, their interactions with neighboring cells, and their responses to the surrounding microenvironment. An important challenge in the field is scaling cellular resolution models to real-scale tumor simulations, which is critical for the development of digital twin models of diseases and requires the use of High-Performance Computing (HPC) since every time step involves trillions of operations. We hereby present a scalable HPC solution for the molecular diffusion modeling using an efficient implementation of state-of-the-art Finite Volume Method (FVM) frameworks. The paper systematically evaluates a novel scalable Biological Finite Volume Method (BioFVM) library and presents an extensive performance analysis of the available solutions. Results shows that our HPC proposal reach almost 200x speedup and up to 36
Synthetic data (SD) has become an increasingly important asset in the life sciences, helping address data scarcity, privacy concerns, and barriers to data access. Creating artificial datasets that mirror the characteristics of real data allows researchers to develop and validate computational methods in controlled environments. Despite its promise, the adoption of SD in life sciences hinges on rigorous evaluation metrics designed to assess their fidelity and reliability. To explore the current landscape of SD evaluation metrics in distinct life sciences domains, the ELIXIR Machine Learning Focus Group performed a systematic review of the scientific literature following the PRISMA guidelines. Six critical domains were examined to identify current practices for assessing SD. Findings reveal that, while generation methods are rapidly evolving, systematic evaluation is often overlooked, limiting researchers' ability to compare, validate, and trust synthetic datasets across different domains. This systematic review underscores the urgent need for robust, standardized evaluation approaches that not only bolster confidence in SD but also guide its effective and responsible implementation. By laying the groundwork for establishing domain-specific yet interoperable standards, this scoping review paves the way for future initiatives aimed at enhancing the role of SD in scientific discovery, clinical practice and beyond.
Summary:Rapid development of genomic technologies in recent years enables personalised medicine to become an essential part of healthcare. Advanced computational methods are required to extract relevant insights that can be applied in clinical settings. This presents a challenge for clinicians and biomedical researchers, who need specialised training to adopt these tools. Within the context of PerMedCoE, the first European Centre of Excellence in Personalised Medicine, we developed and delivered a competency-based training programme to support professionals in the life sciences to work with modelling and simulation tools that integrate omics data to identify biological processes relevant to disease. We identified a set of required competencies in the field and built a series of career profiles with specific competence levels in these. The competencies and profiles contributed to define the focus and target audience of the training activities delivered: a combination of self-paced learning resources, webinars and online and face-to-face synchronous courses. The outputs of the programme (competencies, career profiles and training materials) can be used by biomedical professionals for their own career development or to train others. In addition, the approach can be adopted by other fields with rapid technological advancements and a constant need to upskill professionals. Availability and implementation:The competency framework is reproduced in full in this paper as supplementary material and available on the Competency Hub at https://competency.ebi.ac.uk/framework/permedcoe/2.1.
The Ras protein superfamily comprises small GTPases that share a conserved G-domain but differ in flanking regions, regulation, and cellular roles. Because existing classifications rely mainly on G-domain phylogeny, this superfamily provides a useful test case for assessing whether protein language model embeddings recover biologically meaningful sequence organization consistent with established evolutionary classifications. Here, we analyzed a curated Ras superfamily dataset using three classification schemes: the classical five-family view, a G-domain phylogeny-based classification, and UniProtKB family annotation classification. We compared embeddings from multiple protein language models using supervised classification, unsupervised clustering, and residue-level ablation. Sequence-derived embeddings recovered known Ras superfamily organization across analyses. In supervised analyses, simple linear classifiers achieved high performance, indicating that Ras family and subfamily information is linearly accessible from sequence-derived embeddings. In unsupervised analyses, ESM-C layer 12 gave the strongest full-protein recovery of the a G-domain-based classification, whereas ProstT5 performed best for G-domain embeddings. Residue ablation identified recurrent candidate subfamily-informative regions both within and outside the G-domain, including signals mapping to structurally coherent regions associated with subfamily-specific regulatory or interaction-related functions. Together, these results indicate that protein language model embeddings provide an effective alignment-free representation of Ras family and subfamily organization, and can highlight candidate sequence regions associated with functional specialization from sequence alone.
Constructing multicellular mechanistic models traditionally requires extensive time and computational expertise. We introduce intelligent tool orchestration via Model Context Protocol (MCP) servers, enabling Large Language Model (LLM) agents to act as AI laboratory assistants for rapid model prototyping. We demonstrate this approach by constructing a multiscale model of cancer cell fate in response to TNF using an AI agent connected to MCP servers interfacing with three complementary tools: NeKo for gene regulatory networks construction, MaBoSS for Boolean models simulation, and PhysiCell for setting up multicellular agent-based models. This workflow was executed entirely through natural language interactions, without manual coding, direct parameter editing, or manual modification of generated model files. Through this use case, we identified key principles for biological AI-tool integration, specifically regarding tool granularity, session management, and flexible orchestration. Testing across multiple LLMs demonstrated our framework’s portability, though model-dependent variations emphasize the need for rigorous validation. Ultimately, this work establishes a foundation for AI-assisted rapid prototyping, enabling researchers to explore computational hypotheses more rapidly through natural language interaction.
Biological and medical databases are crucial resources without which biomedical research would come to a halt. However, centralized databases often face vulnerabilities, particularly when dependent on single institutions. In contrast, distributed databases enhance resilience, encourage international collaboration, and promote scientific advancement by supporting shared responsibility and improved accessibility.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark17