Abstract Machine learning methods for protein engineering are rarely interoperable, require bespoke workflows, and remain inaccessible to non-experts. Yet the design problems that matter most – conditional design subject to real-world constraints, multi-objective optimization, and iterative lab-in-the-loop workflows where experimental data continuously refines successive design rounds – demand exactly the kind of flexible, composable infrastructure that no single tool provides. We present evedesign, a unified open-source framework that formalizes conditional biosequence design in a method-agnostic way, enabling complex multiobjective workflows combining supervised and unsupervised models from standardized specifications, and built from the outset to support iterative experimental integration. An interactive web interface facilitates end-to-end design for a broad scientific audience at https://evedesign.bio . We demonstrate evedesign’s utility in antibody engineering, enzyme design, and natural enzyme discovery, and invite open-source community contributions.
Artificial intelligence (AI) is transforming scientific research, including proteomics. In this Perspective, we highlight key mass spectrometry (MS)-based proteomics areas where AI is driving innovation, ranging from protein identification to building AI virtual cells. These include improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and, ultimately, enabling AI virtual cells. Finally, we call for global collaboration among data producers, data consumers and other stakeholders to establish an AI-friendly ecosystem for MS-based proteomics, laying the foundation for transformative advancements in proteomics driven by AI.
cBioPortal for Cancer Genomics is a widely used platform for exploratory, interactive visualization and analysis of large-scale clinico-genomic datasets. cBioPortal provides a range of visualizations and analyses including interactive cohort exploration, OncoPrints, mutation “lollipop” plots, survival analysis, alteration enrichment analysis, and detailed patient-level visualizations. cBioPortal also integrates variant annotations from a variety of sources to facilitate interpretation. The public cBioPortal (https://www.cbioportal.org) is accessed by >40,000 unique visitors each month and hosts data from >460 studies. All data is also available in the cBioPortal Datahub: https://github.com/cBioPortal/datahub. In 2024 we added 76 new studies (∼30,000 samples), including data from the NCI Genomic Data Commons. In addition, >94 instances of cBioPortal are installed at academic institutions and companies worldwide. cBioPortal partners with AACR Project GENIE to provide access to the GENIE cohort in a dedicated instance (https://genie.cbioportal.org). Users can explore the full GENIE cohort of >229,000 clinically sequenced samples from 19 institutions, as well as cohorts with comprehensive clinical annotations including response, outcome, and treatment history, from the GENIE Biopharma Collaborative (BPC). BPC cohorts for NSCLC (∼2,000 samples) and colorectal cancer (∼1,500 samples) are available, with more to come. The past year has brought a variety of enhancements to cBioPortal. A new data type selector on the home page enables users to find studies with specific types of data. The interactive cohort exploration has new ways to explore data with the addition of gene-specific charts to summarize the types of mutations in a gene and the integration of the Plots tab for customizable graphs of any two data attributes. The OncoPrint can now display per group alteration frequency based on any categorical attribute. Variant interpretation is enhanced with the integration of AlphaMissense as a novel annotation source and an update to the latest MutationAssessor data. The patient page also has new visualizations, including mutational signatures and the integration of Chromoscope to visualize structural variations. We also made significant changes to the backend code to improve both the developer and user experience. The backend code was repackaged and upgraded to simplify and improve the development process. In addition, we are working on switching to an Online Analytical Processing (OLAP) database which will bring significant performance improvements. cBioPortal is open source: https://github.com/cBioPortal. Development is a collaborative effort among groups at Memorial Sloan Kettering Cancer Center, Dana-Farber Cancer Institute, Children’s Hospital of Philadelphia, Princess Margaret Cancer Centre, Caris Life Sciences, Bilkent University, SE4BIO and The Hyve. We welcome open source contributions from others in the cancer research community. Ino de Bruijn,Tali Mazor,Rima AlHamad,Calla Chennault,Corey Dubin,Jeremy Easton-Marks,Zhaoyuan Fu,Benjamin Gross,Charles Haynes,David M. Higgins,Jason Hwee,Prasanna K. Jagannathan,Mirella Kalafati,Karthik Kalletla,Zeynep Karagöz,James Ko,Tim Kuijpers,Sowmiyaa Kumar,Priti Kumari,Ritika Kundra,Bryan Lai,Xiang Li,James Lindsay,Aaron Lisman,Qi-Xuan Lu,Ramyasree Madupuri,Zain-ul-Abideen Nasir,Angelica Ochoa,Yusuf Ziya Özgül,Oleguer Plantalech,Matthijs N. Pon,Baby A. Satravada,Jessica Singh,Selcuk Onur Sumer,Pim van Nierop,Floris Vleugels,Avery Wang,Manda Wilson,Hongxin Zhang,Gaofei Zhao,Ugur Dogrusoz,Allison Heath,Adam Resnick,Trevor J. Pugh,Chris Sander,Ethan Cerami,JianJiong Gao,Nikolaus Schultz. cBioPortal for cancer genomics [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 1117.
Abstract cBioPortal is a widely used platform for interactive visualization and analysis of large-scale multimodal cancer datasets. It provides cohort exploration tools, such as OncoPrint, mutation “lollipop” plots, survival and enrichment analyses, detailed patient-level views, and integrated variant annotations to support interpretation. The public instance of cBioPortal (https://www.cbioportal.org) serves >40,000 unique visitors globally each month. It hosts data from >500 studies, all also available through the cBioPortal Datahub. In 2025, we added 38 new studies (∼35,000 samples). We also added Tumor Break Load (TBL) scores across PCAWG, CCLE, and all 32 TCGA Pan-Cancer Atlas studies. More than 99 cBioPortal instances are deployed at institutions and companies worldwide. cBioPortal partners with AACR Project GENIE to provide access to the GENIE cohort in a dedicated instance (https://genie.cbioportal.org). Users can explore >268,000 clinically sequenced samples from 20 institutions, as well as GENIE Biopharma Collaborative (BPC) cohorts with detailed clinical annotations, including NSCLC (∼2,000 samples), colorectal cancer (∼1,500), and breast cancer (∼1,200), with more to come. Over the past year, cBioPortal progressed along two complementary directions. First, we introduced a chat-based interface for natural-language data exploration, reflecting ongoing efforts to utilize AI, specifically large language models, to augment traditional query and visualization workflows in cBioPortal. Second, we released several core platform enhancements: 1) we improved the performance for large cohorts by switching the backend database from MySQL to ClickHouse, an OLAP (Online Analytical Processing) database; 2) the Plots tab can now visualize variant allele frequencies, connect multiple samples from the same patient, and has more flexible categorical sorting; 3) variant interpretation has been strengthened through integration of AlphaMissense predictions; 4) we released a redesigned About page highlighting the year’s accomplishments and future roadmap. It is worth noting that many of these features were developed using AI-assisted technologies, which are increasingly standard practice in software engineering. cBioPortal is open source (https://github.com/cBioPortal) and developed collaboratively by groups at Memorial Sloan Kettering Cancer Center, Dana-Farber Cancer Institute, Children’s Hospital of Philadelphia, Princess Margaret Cancer Centre, Bilkent University, SE4BIO, and The Hyve. We welcome contributions from the cancer research community. Citation Format: Ino de Bruijn, Tali Mazor, Gaofei Zhao, Manda Wilson, Avery Wang, Floris Vleugels, Pim van Nierop, Henk-Jan van den Ham, S. Onur Sumer, Jessica Singh, Baby A. Satravada, Oleguer Plantalech, Angelica Ochoa, Zain-ul-Abideen Nasir, Ramyasree Madupuri, Pieter Lukasse, Aaron Lisman, James Lindsay, Xiang Li, Bryan Lai, Ritika Kundra, Priti Kumari, Sowmiyaa Kumar, Tim Kuijpers, James Ko, Zeynep Karagöz, Karthik Kalletla, Prasanna K Jagannathan, Jason Hwee, Guizela Huelsz Prince, Charles Haynes, Benjamin Gross, Zhaoyuan Fu, Ruslan Forostianov, Calla Chennault, Rima AlHamad, Ugur Dogrusoz, Allison Heath, Adam C. Resnick, Trevor J. Pugh, Chris Sander, Jianjiong Gao, Nikolaus Schultz, Ethan Cerami. cBioPortal for cancer genomics [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 4097.
MutationAssessor (MA) helps researchers evaluate the likely functional impact of somatic and germline mutations in cancer. It provides an evolution-based functional impact score (FIS) to classify mutations based on their likely effect on protein function. FIS scores are based on analysis of patterns of conservation in protein families (conserved residues) and subfamilies (specificity residues). In this new version (r4) we have (1) refined the combinatorial entropy analysis of conservation patterns, (2) recalculated full-length protein multiple sequence alignments covering a larger fraction of human proteins and making use of the explosive growth of protein sequence data, (3) compared predicted functional impact with the pathogenic-benign classification of sequence variants in curated knowledge bases, such as ClinVar, (4) observed the inverse relationship between predicted high functional impact and variant frequency in germline genome sequences and (5) explore the evaluation of switch-of-function mutational effects. Functional impact of ~4 million somatic amino-acid changing mutations across more than 320K human tumor samples are now available in the widely used cBioPortal for Cancer Genomics.
Advances in single-cell technology have enabled the measurement of cell-resolved molecular states across a variety of cell lines and tissues under a plethora of genetic, chemical, environmental or disease perturbations. Current methods focus on differential comparison or are specific to a particular task in a multi-condition setting with purely statistical perspectives. The quickly growing number, size and complexity of such studies require a scalable analysis framework that takes existing biological context into account. Here we present pertpy, a Python-based modular framework for the analysis of large-scale single-cell perturbation experiments. Pertpy provides access to harmonized perturbation datasets and metadata databases along with numerous fast and user-friendly implementations of both established and novel methods, such as automatic metadata annotation or perturbation distances, to efficiently analyze perturbation data. As part of the scverse ecosystem, pertpy interoperates with existing single-cell analysis libraries and is designed to be easily extended.
The relationship between pH and enzyme catalytic activity, especially the optimal pH (pHopt) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pHopt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence–function relationships. Here we proposed and evaluated various machine learning methods for predicting pHopt, conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pHopt, including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pHopt prediction and will potentially speed up the development of enzyme technologies. Accurately predicting the optimal pH level for enzyme activity is challenging due to the complex relationship between enzyme structure and function. Gado and colleagues show that a language model can effectively learn the structural and biophysical features to predict the optimal pH for enzyme activity.
Artificial intelligence (AI) is transforming scientific research, including proteomics. Advances in mass spectrometry (MS)-based proteomics data quality, diversity, and scale, combined with groundbreaking AI techniques, are unlocking new challenges and opportunities in biological discovery. Here, we highlight key areas where AI is driving innovation, from data analysis to new biological insights. These include developing an AI-friendly ecosystem for proteomics data generation, sharing, and analysis; improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and ultimately enabling AI-empowered virtual cells.
The detection of somatic mutagenesis in normal tissues has transformed our understanding of clonal evolution, revealing a mosaic of somatic mutations, including those in cancer driver genes, that influence cellular behavior and tissue homeostasis. However, key questions remain about how these mutational patterns contribute to tumorigenesis and the role of cancer driver mutations in the transition from normal to malignant states. To address this, we performed a combined computational analysis of the largest normal-tissue DNA-sequencing and clinical dataset to date, spanning 21 published datasets, 37 tissue types, 521 patients, and 19,643 samples, and integrated these with precancer and cancer tissue data to identify tissue-specific mutational changes potentially critical for tumorigenesis. We identified tissue-specific differences in mutation burden, with epithelial tissues exhibiting significantly higher mutation rates than other tissue categories (p<2x10-16). These differences were further reflected in the mutational signatures: UV-light signatures dominated sun-exposed epithelial tissues, while gastrointestinal epithelial tissues showed mutation burden driven by intrinsic factors. Regression analysis revealed that, alongside aging, mutations in specific driver genes (e.g., ARID2, ATM, TP53), but not others, were selectively associated with higher mutation rates, highlighting nuanced contributions to tissue mutagenesis. Further MutSig2CV and dN/dS analyses identified positively selected driver genes with distinct patterns based on tumorigenic pathways: tumors that develop through normal-to-dysplastic sequences (e.g., colon cancer) shared driver genes across all progression stages, reflecting clonal evolution within the same tissue lineage. In contrast, tumors arising from metaplastic processes (e.g., esophageal adenocarcinoma) shared mutated driver genes between cancer and precancer metaplastic tissues, but not normal tissues, suggesting selective pressures emerge during metaplasia. Across the cohort, most driver genes were highly mutated in cancer and present at lower mutation frequency in normal tissues. In contrast, similar to NOTCH1 mutations, CHEK2 mutations were significantly enriched in the normal esophagus but absent in cancer, hinting at potential tumor-suppressive effects. Taken together, our large-scale analysis has allowed us to detect previously unrecognized phenomena through increased statistical power. The analysis of normal tissue mutagenesis underscores its potential for predicting tissue-specific cancer risks. In addition, our integration of normal tissue data into the cBioPortal platform provides an important resource to support and advance future cancer research. Leonie Stockschläder, Eldar Abdullaev, Siddharth Annaldasula, Ino de Bruijn, Baby Anusha Satravada, Ritika Kundra, Rima Madupuri, Manuela Benary, Dieter Beule, Tim H. Coorens, Regine Armann, Nikolaus Schultz, Chris Sander, Kirsten Kubler. Comprehensive analysis of somatic mutagenesis in normal tissues across 19,643 samples reveals insights into clonal evolution and tumorigenesis [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3881.
High-grade serous ovarian cancer (HGSOC) remains the most lethal gynecologic malignancy, and novel treatment approaches are needed. Here, we used unbiased quantitative protein mass spectrometry to assess the cellular response profile to drug perturbations in ovarian cancer cells for the rational design of potential combination therapies. Analysis of the perturbation profiles revealed proteins responding across several drug perturbations (called frequently responsive below) as well as drug-specific protein responses. The frequently responsive proteins included proteins that reflected general drug resistance mechanisms, such as changes in drug efflux pumps. Network analysis of drug-specific protein responses revealed known and potential novel markers of resistance, which were used to rationalize the design of anti-resistance drug pairs. We experimentally tested the anti-proliferative effects of 12 of the proposed drug combinations in 6 HGSOC cell lines. While response typically varies across different cell lines, 10 of the 12 combinations tested have either an additive or synergistic CI index in at least one cell line and may therefore be plausible candidates for overcoming or preventing resistance to single agents. Serendipitously, we observed an unexpectedly strong 0.05-0.11 micromolar response to GPX4 inhibitors as single agents in the OVCAR-4 cell line. We propose several drug combinations as potential therapeutic candidates in ovarian cancer, as well as GPX4 inhibitors as single agents.
Pancreatic ductal adenocarcinoma (PDAC) is a rare, aggressive cancer often diagnosed late with low survival rates, due to the lack of population-wide screening programs and the high cost of early detection methods. To enable early detection of high-risk individuals, we develop a transformer-based model trained on longitudinal Veterans Affairs electronic health record (EHR) with 19,426 PDAC cases and ∼15.9 million controls. Our model combines diagnostic and medication trajectories to predict PDAC risk within a 6-, 12-, and 36-month assessment window. Incorporating medication significantly improved performance; among the top 1,000-5,000 highest-risk patients in a cohort of 1 million patients, 3-year PDAC incidence is 115-70 times higher than a reference estimate based on age and sex alone. Furthermore, analysis of most predictive features highlights the role of events such as chronic inflammatory conditions and specific medications on overall PDAC risk. Our work provides an AI-driven identification of high-risk individuals, with a potential to improve early detection, enhance patient care, and reduce healthcare costs.
Analysis across a growing number of single-cell perturbation datasets is hampered by poor data interoperability. To facilitate development and benchmarking of computational methods, we collect a set of 44 publicly available single-cell perturbation–response datasets with molecular readouts, including transcriptomics, proteomics and epigenomics. We apply uniform quality control pipelines and harmonize feature annotations. The resulting information resource, scPerturb, enables development and testing of computational methods, and facilitates comparison and integration across datasets. We describe energy statistics (E-statistics) for quantification of perturbation effects and significance testing, and demonstrate E-distance as a general distance measure between sets of single-cell expression profiles. We illustrate the application of E-statistics for quantifying similarity and efficacy of perturbations. The perturbation–response datasets and E-statistics computation software are publicly available at scperturb.org. This work provides an information resource for researchers working with single-cell perturbation data and recommendations for experimental design, including optimal cell counts and read depth. scPerturb is an information resource for single-cell perturbation data analysis and comparison.
Recent breakthroughs in AI coupled with the rapid accumulation of protein sequence and structure data have radically transformed computational protein design. New methods promise to escape the constraints of natural and laboratory evolution, accelerating the generation of proteins for applications in biotechnology and medicine. To make sense of the exploding diversity of machine learning approaches, we introduce a unifying framework that classifies models on the basis of their use of three core data modalities: sequences, structures and functional labels. We discuss the new capabilities and outstanding challenges for the practical design of enzymes, antibodies, vaccines, nanomachines and more. We then highlight trends shaping the future of this field, from large-scale assays to more robust benchmarks, multimodal foundation models, enhanced sampling strategies and laboratory automation.
The insufficient availability of comprehensive protein-level perturbation data is impeding the widespread adoption of systems biology. In this perspective, we introduce the rationale, essentiality, and practicality of perturbation proteomics. Biological systems are perturbed with diverse biological, chemical, and/or physical factors, followed by proteomic measurements at various levels, including changes in protein expression and turnover, post-translational modifications, protein interactions, transport, and localization, along with phenotypic data. Computational models, employing traditional machine learning or deep learning, identify or predict perturbation responses, mechanisms of action, and protein functions, aiding in therapy selection, compound design, and efficient experiment design. We propose to outline a generic PMMP (perturbation, measurement, modeling to prediction) pipeline and build foundation models or other suitable mathematical models based on large-scale perturbation proteomic data. Finally, we contrast modeling between artificially and naturally perturbed systems and highlight the importance of perturbation proteomics for advancing our understanding and predictive modeling of biological systems.
A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β-lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.
Abstract cBioPortal for Cancer Genomics is an open-source platform for interactive, exploratory analysis of large-scale clinico-genomic data sets. cBioPortal provides a suite of user-friendly visualizations and analyses, including OncoPrints, mutation “lollipop” plots, variant interpretation, group comparison, survival analysis, expression correlation analysis, alteration enrichment analysis, cohort and patient-level visualization. The public site (https://www.cbioportal.org) is accessed by >35,000 unique visitors each month and hosts data from >350 studies spanning individual labs and large consortia. In addition, at least 74 instances of cBioPortal are installed at academic institutions and companies worldwide. To better support all users, we unified our documentation (https://docs.cbioportal.org) and added a user guide and an ongoing series of ‘how-to’ videos to address common questions. In 2022 we added 32 studies (>38,000 samples) to the public site. In addition, we added a nonsynonymous tumor mutation burden (TMB) value for all samples and enhanced the TCGA PanCancer Atlas studies with DNA methylation and treatment data. All data is available in the cBioPortal Datahub: https://github.com/cBioPortal/datahub. We also host a dedicated instance for AACR Project GENIE, enabling access to the GENIE cohort of >165,000 clinically sequenced samples from 19 institutions (https://genie.cbioportal.org). The GENIE Biopharma Collaborative (BPC) enables the collection of comprehensive clinical annotations, including response, outcome, and treatment history. The first BPC cohorts are now available: ~2,000 non-small cell lung cancer samples and ~1,500 colorectal cancer samples. Support for multimodal data analysis has been a major focus, including several new integrations with external tools. Single cell data is now available in the CPTAC GBM study and can be visualized throughout cBioPortal, and via integration with cellxgene. On the patient page, H&E and mIF images can be visualized via integration with Minerva, and the genomic overview now integrates IGV. We continue to enhance existing features. In the study view, users can now add charts comparing categorical vs continuous data, and the plots tab includes a heatmap option. We replaced the existing fusion data type with a generalized structural variant data type that supports detailed information including breakpoints and orientation, to enable new visualizations and analyses. Pathway level analysis has been extended with a new integration with NDEx. cBioPortal is fully open source (https://github.com/cBioPortal/). Development is a collaborative effort among groups at Memorial Sloan Kettering Cancer Center, Dana-Farber Cancer Institute, Children’s Hospital of Philadelphia, Princess Margaret Cancer Centre, Caris Life Sciences, Bilkent University and The Hyve. We welcome open source contributions from others in the cancer research community. Citation Format: Ino de Bruijn, Tali Mazor, Adam Abeshouse, Diana Baiceanu, Stephanie Carrero, Elena Garcia Lara, Benjamin Gross, David M. Higgins, Prasanna K. Jagannathan, Priti Kumari, Ritika Kundra, Bryan Lai, Xiang Li, James Lindsay, Aaron Lisman, Divya Madala, Ramyasree Madupuri, Angelica Ochoa, Yusuf Ziya Özgül, Oleguer Plantalech, Sander Rodenburg, Baby Anusha Satravada, Robert Sheridan, Lucas Sikina, Jessica Singh, S Onur Sumer, Yichao Sun, Pim van Nierop, Avery Wang, Manda Wilson, Hongxin Zhang, Gaofei Zhao, Sjoerd van Hagen, Ugur Dogrusoz, Allison Heath, Adam Resnick, Trevor J. Pugh, Chris Sander, Ethan Cerami, Jianjiong Gao, Nikolaus Schultz. cBioPortal for Cancer Genomics. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 4256.
The human body contains trillions of cells, classified into specific cell types, with diverse morphologies and functions. In addition, cells of the same type can assume different states within an individual's body during their lifetime. Understanding the complexities of the proteome in the context of a human organism and its many potential states is a necessary requirement to understanding human biology, but these complexities can neither be predicted from the genome, nor have they been systematically measurable with available technologies. Recent advances in proteomic technology and computational sciences now provide opportunities to investigate the intricate biology of the human body at unprecedented resolution and scale. Here we introduce a big-science endeavour called π-HuB (proteomic navigator of the human body). The aim of the π-HuB project is to (1) generate and harness multimodality proteomic datasets to enhance our understanding of human biology; (2) facilitate disease risk assessment and diagnosis; (3) uncover new drug targets; (4) optimize appropriate therapeutic strategies; and (5) enable intelligent healthcare, thereby ushering in a new era of proteomics-driven phronesis medicine. This ambitious mission will be implemented by an international collaborative force of multidisciplinary research teams worldwide across academic, industrial and government sectors.