The tumor suppressor TP53 is frequently mutated in hormone receptor-negative, HER2-positive breast cancer (BC), contributing to tumor aggressiveness. Traditional ancillary methods like immunohistochemistry (IHC) to assess TP53 functionality face pre- and post-analytical challenges. This proof-of-concept study employed a deep learning (DL) algorithm to predict TP53 mutational status from H&E-stained whole slide images (WSIs) of BC tissue. Using a pre-trained convolutional neural network, the model identified tumor areas and predicted TP53 mutations with a Dice coefficient score of 0.82. Predictions were validated through IHC and next-generation sequencing (NGS), confirming TP53 aberrant expression in 92% of the tumor area, closely matching IHC findings (90%). The DL model exhibited high accuracy in tissue quantification and TP53 status prediction, outperforming traditional methods in terms of precision and efficiency. DL-based approaches offer significant promise for enhancing biomarker testing and precision oncology by reducing intra- and inter-observer variability, but further validation is required to optimize their integration into real-world clinical workflows. This study underscores the potential of DL algorithms to predict key genetic alterations, such as TP53 mutations, in BC. DL-based histopathological analysis represents a valuable tool for improving patient management and tailoring treatment approaches based on molecular biomarker status.
Motivation Tumor mutational burden (TMB) has been proposed as a predictive biomarker for immunotherapy response in cancer patients, as it is thought to enrich for tumors with high neoantigen load. TMB assessed by whole-exome sequencing is considered the gold standard but remains confined to research settings. In the clinical setting, targeted gene panels sampling various genomic sizes along with diverse strategies to estimate TMB were proposed and no real standard has emerged yet. Results We provide the community with TMBleR, a tool to measure the clinical impact of various strategies of panel-based TMB measurement. Availability and implementation R package and docker container (GPL-3 Open Source license): https://acc-bioinfo.github.io/TMBleR/. Graphical-user interface website: https://bioserver.ieo.it/shiny/app/tmbler. Supplementary information Supplementary data are available at Bioinformatics online.
BACKGROUND:The deeper knowledge of non-small-cell lung cancer (NSCLC) biology and the discovery of driver molecular alterations have opened the era of precision medicine in lung oncology, thus significantly revolutionizing the diagnostic and therapeutic approach to NSCLC. In Italy, however, molecular assessment remains heterogeneous across the country, and numbers of patients accessing personalized treatments remain relatively low. Nationwide programs have demonstrated that the creation of consortia represent a successful strategy to increase the number of patients with a molecular classification.PATIENTS AND METHODS:The Alliance Against Cancer (ACC), a network of 25 Italian Research Institutes, has developed a targeted sequencing panel for the detection of genomic alterations in 182 genes in patients with a diagnosis of NSCLC (ACC lung panel). One thousand metastatic NSCLC patients will be enrolled onto a prospective trial designed to measure the sensitivity and specificity of the ACC lung panel as a tool for molecular screening compared to standard methods.RESULTS AND CONCLUSION:The ongoing trial is part of a nationwide strategy of ACC to develop infrastructures and improve competences to make the Italian research institutes independent for genomic profiling of cancer patients.
The 17th International NETTAB workshop was held in Palermo, Italy, on October 16-18, 2017. The special topic for the meeting was "Methods, tools and platforms for Personalised Medicine in the Big Data Era", but the traditional topics of the meeting series were also included in the event. About 40 scientific contributions were presented, including four keynote lectures, five guest lectures, and many oral communications and posters. Also, three tutorials were organised before and after the workshop. Full papers from some of the best works presented in Palermo were submitted for this Supplement of BMC Bioinformatics. Here, we provide an overview of meeting aims and scope. We also shortly introduce selected papers that have been accepted for publication in this Supplement, for a complete presentation of the outcomes of the meeting.
In NSCLC, large-scale mutational analysis facilitates access to targeted treatments but is still not routinely employed due to significant technological barriers. The Alleanza Contro il Cancro (ACC) network of Italian Cancer Centers developed an affordable targeted sequencing panel for the identification of multiple genetic alterations with potential clinical utility, and designed a prospective multicentric trial to recruit 1000 newly diagnosed advanced NSCLC patients, aiming to i) compare panel performance against a set of externally validated biomarkers, including alterations in standard-of-care (EGFR, ROS1 and ALK) and non-standard-of-care (KRAS, BRAF, MET) biomarkers; ii) identify alterations in a large dataset of driver and potentially actionable genes; iii) correlate genotypes to survival outcomes and toxicity; iv) carry out ancillary studies on additional biomarkers and/or on specific patient groups (e.g. mutational burden, cfDNA, extensive characterization of immunotherapy-treated patients); v) build a centralized data repository for mutation interpretation and clinical recommendation. Through systematic literature mining and ad-hoc developed bioinformatic pipelines we identified: i) a set of 164 potentially actionable genes in solid tumors; ii) additional 18 genes with predicted driver function in NSCLC; iii) 70 actionable fusion transcripts; iii) 141 SNPs associated with pharmacogenomics markers. We designed a custom enrichment panel (∼800 kb target) and compared PCR- and hybridization-based enrichment on semiconductor or by-synthesis sequencing to be subsequently deployed in a large observational trial. Sequencing is decentralized, to allow rapid turnaround time, but raw and processed data are collected in a single informatic infrastructure for centralized quality control and continuous bioinformatic pipeline improvement. PCR/semiconductor sequencing was selected for deployment based on cost and feasibility (2-day, highly automated workflow). 182 patients have been enrolled to date (90% stage IV, 10% IIIB). Of 65 patients with treatment information available, 28 (43%) subsequently received immunotherapy and 13 (20%) targeted therapy. For 56 patients with complete sequencing data, EGFR and KRAS status was concordant in 9/10 and 38/41 cases; discordant cases are being validated with orthogonal methods. Clinically significant MET amplifications were called in 2/2 cases. Remaining target regions did not show pathogenic alterations. Multiple alterations in potentially actionable genes were identified. Large-scale sequencing is reliable, feasible and sustainable across multiple hospitals and provides clinically relevant results. The increased availability of genomic information may result in enhanced access to tailored therapies. Data and sample integration in centralized, shared repositories will allow multiple ancillary studies.
BACKGROUND:Systems biologists study interaction data to understand the behaviour of whole cell systems, and their environment, at a molecular level. In order to effectively achieve this goal, it is critical that researchers have high quality interaction datasets available to them, in a standard data format, and also a suite of tools with which to analyse such data and form experimentally testable hypotheses from them. The PSI-MI XML standard interchange format was initially published in 2004, and expanded in 2007 to enable the download and interchange of molecular interaction data. PSI-XML2.5 was designed to describe experimental data and to date has fulfilled this basic requirement. However, new use cases have arisen that the format cannot properly accommodate. These include data abstracted from more than one publication such as allosteric/cooperative interactions and protein complexes, dynamic interactions and the need to link kinetic and affinity data to specific mutational changes.RESULTS:The Molecular Interaction workgroup of the HUPO-PSI has extended the existing, well-used XML interchange format for molecular interaction data to meet new use cases and enable the capture of new data types, following extensive community consultation. PSI-MI XML3.0 expands the capabilities of the format beyond simple experimental data, with a concomitant update of the tool suite which serves this format. The format has been implemented by key data producers such as the International Molecular Exchange (IMEx) Consortium of protein interaction databases and the Complex Portal.CONCLUSIONS:PSI-MI XML3.0 has been developed by the data producers, data users, tool developers and database providers who constitute the PSI-MI workgroup. This group now actively supports PSI-MI XML2.5 as the main interchange format for experimental data, PSI-MI XML3.0 which additionally handles more complex data types, and the simpler, tab-delimited MITAB2.5, 2.6 and 2.7 for rapid parsing and download.
Background The increasing availability of resequencing data has led to a better understanding of the most important genes in cancer development. Nevertheless, the mutational landscape of many tumor types is heterogeneous and encompasses a long tail of potential driver genes that are systematically excluded by currently available methods due to the low frequency of their mutations. We developed LowMACA (Low frequency Mutations Analysis via Consensus Alignment), a method that combines the mutations of various proteins sharing the same functional domains to identify conserved residues that harbor clustered mutations in multiple sequence alignments. LowMACA is designed to visualize and statistically assess potential driver genes through the identification of their mutational hotspots. Results We analyzed the Ras superfamily exploiting the known driver mutations of the trio K-N-HRAS, identifying new putative driver mutations and genes belonging to less known members of the Rho, Rab and Rheb subfamilies. Furthermore, we applied the same concept to a list of known and candidate driver genes, and observed that low confidence genes show similar patterns of mutation compared to high confidence genes of the same protein family. Conclusions LowMACA is a software for the identification of gain-of-function mutations in putative oncogenic families, increasing the amount of information on functional domains and their possible role in cancer. In this context LowMACA emphasizes the role of genes mutated at low frequency otherwise undetectable by classical single gene analysis. LowMACA is an R package available at http://www.bioconductor.org/packages/release/bioc/html/LowMACA.html . It is also available as a GUI standalone downloadable at: https://cgsb.genomics.iit.it/wiki/projects/LowMACA
Background Biologists generally interrogate genomics data using web-based genome browsers that have limited analytical potential. New generation genome browsers such as the Integrated Genome Browser (IGB) have largely overcome this limitation and permit customized analyses to be implemented using plugins. We illustrate the use of a plugin for IGB that exploits advanced visualization techniques to integrate the analysis of genomics data with network and structural approaches. Results We show how visualization technologies that combine both genomics and network biology can facilitate the selection of the key amino acid contacts from protein-protein and protein-drug interactions. Starting from the MDM2-P53 interaction, which is a high-value target for cancer therapy, and Nutlin, the parent small molecule of an MDM2 antagonist that is currently in clinical trials, we show that this method can be generalized to analyze how drugs and mutations can interfere with both protein-protein and drug-protein networks. We illustrate this point by two additional use-cases exploring the molecular basis of tamoxifen side effects and of drug resistance in chronic myeloid leukemia patients. Conclusions Combined network and structure biology approaches provide key insights into both the genetic and the edgetic roles of variants in diseases. 3D interactomes facilitate the identification of disease-relevant interactions that can then be specifically targeted by drugs. Recent advances in molecular interaction and structure visualization tools have greatly simplified the mapping of mutated residues to molecular interaction interfaces. Such approaches can now also be integrated with genome visualization tools to enable comparative analyses of interaction contacts.
Motivation Next Generation Sequencing (NGS), with the high amount of heterogeneous data that it is generating, is opening many interesting practical and theoretical computational problems. Genome browsers, e.g. UCSC Genome Browser (Kuhn et al., 2013) or Integrated Genome Browser (IGB) (Nicol et al., 2009), allow visual inspection and identification of interesting patterns on multiple genome browser tracks, i.e. of sets of (epi)genomic regions/peaks at given distances from each other in different tracks. For example, such patterns can describe gene expression regulatory DNA areas including heterogeneous (epi)genomic features (e.g. histone modification and/or different transcription factor binding regions). Yet, once such patterns are visually identified in a genome section, the search of their occurrences along the whole genome is a complex computational task that is currently not supported, despite their discovery along the whole genome is very important for the biological interpretation of NGS experimental results and comprehension of biomolecular phenomena. We defined an optimized pattern-search algorithm able to find efficiently, within a large set of (epi)genomic data, genomic region sets which are similar to a given pattern. We implemented it within an IGB plugin, which allows intuitive user interaction in both the visual selection of an interesting pattern on the loaded IGB tracks, and the visualization of occurrences of similar patterns identified along the entire genome.
Next-generation sequencing (NGS) technologies have deeply changed our understanding of cellular processes by delivering an astonishing amount of data at affordable prices; nowadays, many biology laboratories have already accumulated a large number of sequenced samples. However, managing and analyzing these data poses new challenges, which may easily be underestimated by research groups devoid of IT and quantitative skills. In this perspective, we identify five issues that should be carefully addressed by research groups approaching NGS technologies. In particular, the five key issues to be considered concern: (1) adopting a laboratory management system (LIMS) and safeguard the resulting raw data structure in downstream analyses; (2) monitoring the flow of the data and standardizing input and output directories and file names, even when multiple analysis protocols are used on the same data; (3) ensuring complete traceability of the analysis performed; (4) enabling non-experienced users to run analyses through a graphical user interface (GUI) acting as a front-end for the pipelines; (5) relying on standard metadata to annotate the datasets, and when possible using controlled vocabularies, ideally derived from biomedical ontologies. Finally, we discuss the currently available tools in the light of these issues, and we introduce HTS-flow, a new workflow management system conceived to address the concerns we raised. HTS-flow is able to retrieve information from a LIMS database, manages data analyses through a simple GUI, outputs data in standard locations and allows the complete traceability of datasets, accompanying metadata and analysis scripts.
Summary: Prioritization of candidate genes emanating from large-scale screens requires integrated analyses at the genomics, molecular, network and structural biology levels. We have extended the Integrated Genome Browser (IGB) to facilitate these tasks. The graphical user interface greatly simplifies building disease networks and zooming in at atomic resolution to identify variations in molecular complexes that may affect molecular interactions in the context of genomic data. All results are summarized in genome tracks and can be visualized and analyzed at the transcript level. Availability and implementation: The MI Bundle is a plugin for the IGB. The plugin, help, video and tutorial are available at http://cru.genomics.iit.it/igbmibundle/ and https://github.com/CRUiit/igb-mi-bundle/wiki. The source code is released under the Apache License, Version 2. Contact: arnaud.ceol@iit.it Supplementary information: Supplementary data are available at Bioinformatics online.
Helicobacter pylori infections cause gastric ulcers and play a major role in the development of gastric cancer. In 2001, the first protein interactome was published for this species, revealing over 1500 binary protein interactions resulting from 261 yeast two-hybrid screens. Here we roughly double the number of previously published interactions using an ORFeome-based, proteome-wide yeast two-hybrid screening strategy. We identified a total of 1515 protein-protein interactions, of which 1461 are new. The integration of all the interactions reported in H. pylori results in 3004 unique interactions that connect about 70% of its proteome. Excluding interactions of promiscuous proteins we derived from our new data a core network consisting of 908 interactions. We compared our data set to several other bacterial interactomes and experimentally benchmarked the conservation of interactions using 365 protein pairs (interologs) of E. coli of which one third turned out to be conserved in both species.
Enterohemorrhagic E. coli (EHEC) manipulate their human host through at least 39 effector proteins which hijack host processes through direct protein-protein interactions (PPIs). To identify their protein targets in the host cells, we performed yeast two-hybrid screens, allowing us to find 48 high-confidence protein-protein interactions between 15 EHEC effectors and 47 human host proteins. In comparison to other bacteria and viruses we found that EHEC effectors bind more frequently to hub proteins as well as to proteins that participate in a higher number of protein complexes. The data set includes six new interactions that involve the translocated intimin receptor (TIR), namely HPCAL1, HPCAL4, NCALD, ARRB1, PDE6D, and STK16. We compared these TIR interactions in EHEC and enteropathogenic E. coli (EPEC) and found that five interactions were conserved. Notably, the conserved interactions included those of serine/threonine kinase 16 (STK16), hippocalcin-like 1 (HPCAL1) as well as neurocalcin-delta (NCALD). These proteins co-localize with the infection sites of EPEC. Furthermore, our results suggest putative functions of poorly characterized effectors (EspJ, EspY1). In particular, we observed that EspJ is connected to the microtubule system while EspY1 appears to be involved in apoptosis/cell cycle regulation.
Yeast-two hybrid screening of E. coli proteins and integration with protein structure and genetic interaction data provides an extensive interactome resource. Efforts to map the Escherichia coli interactome have identified several hundred macromolecular complexes, but direct binary protein-protein interactions (PPIs) have not been surveyed on a large scale. Here we performed yeast two-hybrid screens of 3,305 baits against 3,606 preys (∼70% of the E. coli proteome) in duplicate to generate a map of 2,234 interactions, which approximately doubles the number of known binary PPIs in E. coli. Integration of binary PPI and genetic-interaction data revealed functional dependencies among components involved in cellular processes, including envelope integrity, flagellum assembly and protein quality control. Many of the binary interactions that we could map in multiprotein complexes were informative regarding internal topology of complexes and indicated that interactions in complexes are substantially more conserved than those interactions connecting different complexes. This resource will be useful for inferring bacterial gene function and provides a draft reference of the basic physical wiring network of this evolutionarily important model microbe.
BACKGROUND:Life-science laboratories make increasing use of Next Generation Sequencing (NGS) for studying bio-macromolecules and their interactions. Array-based methods for measuring gene expression or protein-DNA interactions are being replaced by RNA-Seq and ChIP-Seq. Sequencing is generally performed by specialized facilities that have to keep track of sequencing requests, trace samples, ensure quality and make data available according to predefined privileges. An integrated tool helps to troubleshoot problems, to maintain a high quality standard, to reduce time and costs. Commercial and non-commercial tools called LIMS (Laboratory Information Management Systems) are available for this purpose. However, they often come at prohibitive cost and/or lack the flexibility and scalability needed to adjust seamlessly to the frequently changing protocols employed. In order to manage the flow of sequencing data produced at the Genomic Unit of the Italian Institute of Technology (IIT), we developed SMITH (Sequencing Machine Information Tracking and Handling).METHODS:SMITH is a web application with a MySQL server at the backend. Wet-lab scientists of the Centre for Genomic Science and database experts from the Politecnico of Milan in the context of a Genomic Data Model Project developed SMITH. The data base schema stores all the information of an NGS experiment, including the descriptions of all protocols and algorithms used in the process. Notably, an attribute-value table allows associating an unconstrained textual description to each sample and all the data produced afterwards. This method permits the creation of metadata that can be used to search the database for specific files as well as for statistical analyses.RESULTS:SMITH runs automatically and limits direct human interaction mainly to administrative tasks. SMITH data-delivery procedures were standardized making it easier for biologists and analysts to navigate the data. Automation also helps saving time. The workflows are available through an API provided by the workflow management system. The parameters and input data are passed to the workflow engine that performs de-multiplexing, quality control, alignments, etc.CONCLUSIONS:SMITH standardizes, automates, and speeds up sequencing workflows. Annotation of data with key-value pairs facilitates meta-analysis.
BACKGROUND:Modern genomic technologies produce large amounts of data that can be mapped to specific regions in the genome. Among the first steps in interpreting the results is annotation of genomic regions with known features such as genes, promoters, CpG islands etc. Several tools have been published to perform this task. However, using these tools often requires a significant amount of bioinformatics skills and/or downloading and installing dedicated software.RESULTS:Here we present AnnotateGenomicRegions, a web application that accepts genomic regions as input and outputs a selection of overlapping and/or neighboring genome annotations. Supported organisms include human (hg18, hg19), mouse (mm8, mm9, mm10), zebrafish (danRer7), and Saccharomyces cerevisiae (sacCer2, sacCer3). AnnotateGenomicRegions is accessible online on a public server or can be installed locally. Some frequently used annotations and genomes are embedded in the application while custom annotations may be added by the user.CONCLUSIONS:The increasing spread of genomic technologies generates the need for a simple-to-use annotation tool for genomic regions that can be used by biologists and bioinformaticians alike. AnnotateGenomicRegions meets this demand. AnnotateGenomicRegions is an open-source web application that can be installed on any personal computer or institute server. AnnotateGenomicRegions is available at: http://cru.genomics.iit.it/AnnotateGenomicRegions.