Structural variant (SV) detection in human genomes using short-read sequencing data is hindered by false positives, arising from sequencing and mapping artifacts that mimic genuine SV signals. Despite advances, state-of-the-art SV callers like GRIDSS and Manta exhibit trade-offs between precision and recall, with GRIDSS offering the highest precision and Manta excelling in recall. To address these limitations, we introduce sv-channels, a novel deep learning model designed to improve the precision of SV detection by leveraging read information at call sites. Our method effectively reduces false positives in Manta's deletion callsets, achieving precision that surpasses GRIDSS while maintaining a recall rate comparable to Manta. This represents a significant improvement in SV detection, leveraging Manta's high recall through deep learning and paving the way for more accurate genomic analyses. The sv-channels codebase is openly accessible on GitHub at https://github.com/GooglingTheCancerGenome/sv-channels enabling further research and application in the field. ### Competing Interest Statement J.d.R and W.P.K are co-founders and directors of Cyclomics, a genomics company, they declare no competing interests. L.S is an employer of JSR Life Sciences, he declares no competing interests. S.G, A.K, B.S.P, C.S, S.M, L.R declare no competing interests.
Smoking tobacco and physical inactivity are key preventable behavioural risk factors of cardiovascular disease (CVD). Computerised coaching systems can help individuals to modify risky behaviours, thereby preventing CVD. However, most reported eHealth or computerized coaching systems are hard to reuse in slightly different settings. To provide an open-source, reusable computer coaching system, we developed Perfect Fit. The reusability is manifested by building around the open-source text- and voice-based contextual assistant framework Rasa. Rasa provides a simple, standard interface to many popular messaging and voice channels, and custom connectors are easily implemented. A set of algorithms have been developed and connected to Rasa to drive and personalize the conversation flow and the coaching process. Such algorithms make use of data stored in a devoted database. Furthermore, Perfect Fit adheres to best practices and standards in software engineering. The modular design of Perfect Fit will allow researchers to connect the virtual coach to any messaging or voice channel with only modest modification. Perfect Fit is available under open-source license in GitHub and is currently in prototype-phase. Concluding, Perfect Fit will deliver a virtual coach that can easily be adapted and reused in different settings. The coach helps individuals to achieve and maintain abstinence from smoking and sufficient physical activity (PA).
Mass spectrometry data is one of the key sources of information in many workflows in medicine and across the life sciences. Mass fragmentation spectra are considered characteristic signatures of the chemical compound they originate from, yet the chemical structure itself usually cannot be easily deduced from the spectrum. Often, spectral similarity measures are used as a proxy for structural similarity but this approach is strongly limited by a generally poor correlation between both metrics. Here, we propose MS2DeepScore: a novel Siamese neural network to predict the structural similarity between two chemical structures solely based on their MS/MS fragmentation spectra. Using a cleaned dataset of >100,000 mass spectra of about 15,000 unique known compounds, MS2DeepScore learns to predict structural similarity scores for spectrum pairs with high accuracy. In addition, sampling different model varieties through Monte-Carlo Dropout is used to further improve the predictions and assess the model’s prediction uncertainty. On 3,600 spectra of 500 unseen compounds, MS2DeepScore is able to identify highly-reliable structural matches and predicts Tanimoto scores with a root mean squared error of about 0.15. The prediction uncertainty estimate can be used to select a subset of predictions with a root mean squared error of about 0.1. We demonstrate that MS2DeepScore outperforms classical spectral similarity measures in retrieving chemically related compound pairs from large mass spectral datasets, thereby illustrating its potential for spectral library matching. Finally, MS2DeepScore can also be used to create chemically meaningful mass spectral embeddings that could be used to cluster large numbers of spectra. Added to the recently introduced unsupervised Spec2Vec metric, we believe that machine learning-supported mass spectral similarity metrics have great potential for a range of metabolomics data processing pipelines. ### Competing Interest Statement The authors have declared no competing interest.
Spectral similarity is used as a proxy for structural similarity in many tandem mass spectrometry (MS/MS) based metabolomics analyses such as library matching and molecular networking. Although weaknesses in the relationship between spectral similarity scores and the true structural similarities have been described, little development of alternative scores has been undertaken. Here, we introduce Spec2Vec, a novel spectral similarity score inspired by a natural language processing algorithm -- Word2Vec. Spec2Vec learns fragmental relationships within a large set of spectral data to derive abstract spectral embeddings that can be used to assess spectral similarities. Using data derived from GNPS MS/MS libraries including spectra for nearly 13,000 unique molecules, we show how Spec2Vec scores correlate better with structural similarity than cosine-based scores. We demonstrate the advantages of Spec2Vec in library matching and molecular networking. Spec2Vec is computationally more scalable allowing structural analogue searches in large databases within seconds.
Accurate and low-cost sleep measurement tools are needed in both clinical and epidemiological research. To this end, wearable accelerometers are widely used as they are both low in price and provide reasonably accurate estimates of movement. Techniques to classify sleep from the high-resolution accelerometer data primarily rely on heuristic algorithms. In this paper, we explore the potential of detecting sleep using Random forests. Models were trained using data from three different studies where 134 adult participants (70 with sleep disorder and 64 good healthy sleepers) wore an accelerometer on their wrist during a one-night polysomnography recording in the clinic. The Random forests were able to distinguish sleep-wake states with an F1 score of 73.93% on a previously unseen test set of 24 participants. Detecting when the accelerometer is not worn was also successful using machine learning ([Formula: see text]), and when combined with our sleep detection models on day-time data provide a sleep estimate that is correlated with self-reported habitual nap behaviour ([Formula: see text]). These Random forest models have been made open-source to aid further research. In line with literature, sleep stage classification turned out to be difficult using only accelerometer data.
Mass spectrometry data is one of the key sources of information in many workflows in medicine and across the life sciences. Mass fragmentation spectra are generally considered to be characteristic signatures of the chemical compound they originate from, yet the chemical structure itself usually cannot be easily deduced from the spectrum. Often, spectral similarity measures are used as a proxy for structural similarity but this approach is strongly limited by a generally poor correlation between both metrics. Here, we propose MS2DeepScore: a novel Siamese neural network to predict the structural similarity between two chemical structures solely based on their MS/MS fragmentation spectra. Using a cleaned dataset of > 100,000 mass spectra of about 15,000 unique known compounds, we trained MS2DeepScore to predict structural similarity scores for spectrum pairs with high accuracy. In addition, sampling different model varieties through Monte-Carlo Dropout is used to further improve the predictions and assess the model's prediction uncertainty. On 3600 spectra of 500 unseen compounds, MS2DeepScore is able to identify highly-reliable structural matches and to predict Tanimoto scores for pairs of molecules based on their fragment spectra with a root mean squared error of about 0.15. Furthermore, the prediction uncertainty estimate can be used to select a subset of predictions with a root mean squared error of about 0.1. Furthermore, we demonstrate that MS2DeepScore outperforms classical spectral similarity measures in retrieving chemically related compound pairs from large mass spectral datasets, thereby illustrating its potential for spectral library matching. Finally, MS2DeepScore can also be used to create chemically meaningful mass spectral embeddings that could be used to cluster large numbers of spectra. Added to the recently introduced unsupervised Spec2Vec metric, we believe that machine learning-supported mass spectral similarity measures have great potential for a range of metabolomics data processing pipelines.
Three-dimensional (3D) structures of protein complexes provide fundamental information to decipher biological processes at the molecular scale. The vast amount of experimentally and computationally resolved protein-protein interfaces (PPIs) offers the possibility of training deep learning models to aid the predictions of their biological relevance. We present here DeepRank, a general, configurable deep learning framework for data mining PPIs using 3D convolutional neural networks (CNNs). DeepRank maps features of PPIs onto 3D grids and trains a user-specified CNN on these 3D grids. DeepRank allows for efficient training of 3D CNNs with data sets containing millions of PPIs and supports both classification and regression. We demonstrate the performance of DeepRank on two distinct challenges: The classification of biological versus crystallographic PPIs, and the ranking of docking models. For both problems DeepRank is competitive with, or outperforms, state-of-the-art methods, demonstrating the versatility of the framework for research in structural biology.
Genomics and metabolomics are widely used to explore specialized metabolite diversity. The Paired Omics Data Platform is a community initiative to systematically document links between metabolome and (meta)genome data, aiding identification of natural product biosynthetic origins and metabolite structures.
There has been a large focus in recent years on making assets in scientific research findable, accessible, interoperable and reusable, collectively known as the FAIR principles. A particular area of focus lies in applying these principles to scientific computational workflows. Jupyter notebooks are a very popular medium by which to program and communicate computational scientific analyses. However, they present unique challenges when it comes to reuse of only particular steps of an analysis without disrupting the usual flow and benefits of the notebook approach, making it difficult to fully comply with the FAIR principles. Here we present an approach and toolset for adding the power of semantic technologies to Python-encoded scientific workflows in a simple, automated and minimally intrusive manner. The semantic descriptions are published as a series of nanopublications that can be searched and used in other notebooks by means of a Jupyter Lab plugin. We describe the implementation of the proposed approach and toolset, and provide the results of a user study with 15 participants, designed around image processing workflows, to evaluate the usability of the system and its perceived effect on FAIRness. Our results show that our approach is feasible and perceived as user-friendly. Our system received an overall score of 78.75 on the System Usability Scale, which is above the average score reported in the literature.
It is essential for the advancement of science that researchers share, reuse and reproduce each other’s workflows and protocols. The FAIR principles are a set of guidelines that aim to maximize the value and usefulness of research data, and emphasize the importance of making digital objects findable and reusable by others. The question of how to apply these principles not just to data but also to the workflows and protocols that consume and produce them is still under debate and poses a number of challenges. In this paper we describe a two-fold approach of simultaneously applying the FAIR principles to scientific workflows as well as the involved data. We apply and evaluate our approach on the case of the PREDICT workflow, a highly cited drug repurposing workflow. This includes FAIRification of the involved datasets, as well as applying semantic technologies to represent and store data about the detailed versions of the general protocol, of the concrete workflow instructions, and of their execution traces. We propose a semantic model to address these specific requirements and was evaluated by answering competency questions. This semantic model consists of classes and relations from a number of existing ontologies, including Workflow4ever, PROV, EDAM, and BPMN. This allowed us then to formulate and answer new kinds of competency questions. Our evaluation shows the high degree to which our FAIRified OpenPREDICT workflow now adheres to the FAIR principles and the practicality and usefulness of being able to answer our new competency questions.
We present the QMflows Python package for quantum chemistry workflow automatization. QMflows allows users to write complex workflows in terms of simple Python scripts. It supports the development of interoperable workflows involving multiple quantum chemistry codes and executes them efficiently on large scale parallel computers. This open source library provides standardized interfaces to a number of quantum chemistry packages and can be easily extended to accommodate additional codes. QMflows features are described and illustrated with a number of representative applications.
This report describes a density functional theory investigation into the reactivities of a series of aza-1,3-dipoles with ethylene at the BP86/TZ2P level. A benchmark study was carried out using QMflows, a newly developed program for automated workflows of quantum chemical calculations. In total, 24 1,3-dipolar cycloaddition (1,3-DCA) reactions were benchmarked using the highly accurate G3B3 method as a reference. We screened a number of exchange and correlation functionals, including PBE, OLYP, BP86, BLYP, both with and without explicit dispersion corrections, to assess their accuracies and to determine which of these computationally efficient functionals performed the best for calculating the energetics for cycloaddition reactions. The BP86/TZ2P method produced the smallest errors for the activation and reaction enthalpies. Then, to understand the factors controlling the reactivity in these reactions, seven archetypal aza-1,3-dipolar cycloadditions were investigated using the activation strain model and energy decomposition analysis. Our investigations highlight the fact that differences in activation barrier for these 1,3-DCA reactions do not arise from differences in strain energy of the dipole, as previously proposed. Instead, relative reactivities originate from differences in interaction energy. Analysis of the 1,3-dipole–dipolarophile interactions reveals the reactivity trends primarily result from differences in the extent of the primary orbital interactions.
This report describes a density functional theory investigation into the reactivities of a series of aza-1,3-dipoles with ethylene at the BP86/TZ2P level. A benchmark study was carried out using QMflows, a newly developed program for automated workflows of quantum chemical calculations. In total, 24 1,3-dipolar cycloaddition (1,3-DCA) reactions were benchmarked using the highly accurate G3B3 method as a reference. We screened a number of exchange and correlation functionals, including PBE, OLYP, BP86, BLYP, both with and without explicit dispersion corrections, to assess their accuracies and to determine which of these computationally efficient functionals performed the best for calculating the energetics for cycloaddition
Complex metabolite mixtures are challenging to unravel. Mass spectrometry (MS) is a widely used and sensitive technique for obtaining structural information of complex mixtures. However, just knowing the molecular masses of the mixture's constituents is almost always insufficient for confident assignment of the associated chemical structures. Structural information can be augmented through MS fragmentation experiments whereby detected metabolites are fragmented, giving rise to MS/MS spectra. However, how can we maximize the structural information we gain from fragmentation spectra? We recently proposed a substructure-based strategy to enhance metabolite annotation for complex mixtures by considering metabolites as the sum of (bio)chemically relevant moieties that we can detect through mass spectrometry fragmentation approaches. Our MS2LDA tool allows us to discover - unsupervised - groups of mass fragments and/or neutral losses, termed Mass2Motifs, that often correspond to substructures. After manual annotation, these Mass2Motifs can be used in subsequent MS2LDA analyses of new datasets, thereby providing structural annotations for many molecules that are not present in spectral databases. Here, we describe how additional strategies, taking advantage of (i) combinatorial in silico matching of experimental mass features to substructures of candidate molecules, and (ii) automated machine learning classification of molecules, can facilitate semi-automated annotation of substructures. We show how our approach accelerates the Mass2Motif annotation process and therefore broadens the chemical space spanned by characterized motifs. Our machine learning model used to classify fragmentation spectra learns the relationships between fragment spectra and chemical features. Classification prediction on these features can be aggregated for all molecules that contribute to a particular Mass2Motif and guide Mass2Motif annotations. To make annotated Mass2Motifs available to the community, we also present MotifDB: an open database of Mass2Motifs that can be browsed and accessed programmatically through an Application Programming Interface (API). MotifDB is integrated within ms2lda.org, allowing users to efficiently search for characterized motifs in their own experiments. We expect that with an increasing number of Mass2Motif annotations available through a growing database, we can more quickly gain insight into the constituents of complex mixtures. This will allow prioritization towards novel or unexpected chemistries and faster recognition of known biochemical building blocks.
Epilepsy is largely under-diagnosed in low-income and middle-income countries, due to lack of medical specialists and expensive electroencephalography (EEG) hardware. In this study we investigate if low-cost consumer-grade EEG in combination with machine learning techniques can offer a reliable screening tool to improve diagnosis rates.We acquired brain signals in people with epilepsy (N=163) and healthy controls (N=138) in two difficult-to-reach areas in rural Guinea-Bissau and Nigeria. Five minutes of fourteen channel resting-state EEG data were acquired with a portable, low-cost consumer-grade EEG recording headset. EEG channel time-series were divided in four-second artifact-free epochs and transformed into delta, theta, alpha, beta and gamma wavelet frequencies. Summary measures such as the mean, standard deviation, minimal value and maximal value of the epoch signal fluctuations were used to train a random forest classifier. Epilepsy diagnosis based on at least three months seizure calendar data was used as the gold standard diagnosis. To prevent too optimistic classification the trained model was evaluated with EEG data from subjects not used in the training. In addition, we tested a classification model trained on Nigeria data against data from people in Guinea-Bissau and vice versa. The most contributing data features in the EEG were found in the beta and theta frequencies in Guinea-Bissau and Nigeria, respectively. Within-country model performance was good with area under the receiver-operating curves of 0.85 and 0.78 (± 0.02 standard errors) in unseen data in Guinea-Bissau and Nigeria, respectively. Across-country performance was moderate (0.62 and 0.64 ± 0.02).Our data suggests that a combination of low cost electroencephalography and machine learning techniques may facilitate diagnostic screening for epilepsy in the most remote areas of the world.
3D-e-Chem-VM is an open source, freely available Virtual Machine (http://3d-e-chem.github.io/3D-e-Chem-VM/) that integrates cheminformatics and bioinformatics tools for the analysis of protein–ligand interaction data. 3D-e-Chem-VM consists of software libraries, and database and workflow tools that can analyze and combine small molecule and protein structural information in a graphical programming environment. New chemical and biological data analytics tools and workflows have been developed for the efficient exploitation of structural and pharmacological protein–ligand interaction data from proteomewide databases (e.g., ChEMBLdb and PDB), as well as customized information systems focused on, e.g., G protein-coupled receptors (GPCRdb) and protein kinases (KLIFS). The integrated structural cheminformatics research infrastructure compiled in the 3D-e-Chem-VM enables the design of new approaches in virtual ligand screening (Chemdb4VS), ligand-based metabolism prediction (SyGMa), and structure-based protein binding site comparison and bioisosteric replacement for ligand design (KRIPOdb).
New Findings What is the central question of this study? Exercise is known to induce stress-related physiological responses, such as changes in intestinal barrier function. Our aim was to determine the test-retest repeatability of these responses in well-trained individuals.What is the main finding and its importance? Responses to strenuous exercise, as indicated by stress-related markers such as intestinal integrity markers and myokines, showed high test-retest variation. Even in well-trained young men an adapted response is seen after a single repetition after 1 week. This finding has implications for the design of studies aimed at evaluating physiological responses to exercise.Strenuous exercise induces different stress-related physiological changes, potentially including changes in intestinal barrier function. In the Protege Study (ISRCTN14236739; ), we determined the test-retest repeatability in responses to exercise in well-trained individuals. Eleven well-trained men (27 +/- 4 years old) completed an exercise protocol that consisted of intensive cycling intervals, followed by an overnight fast and an additional 90min cycling phase at 50% of maximal workload the next morning. The day before (rest), and immediately after the exercise protocol (exercise) a lactulose and rhamnose solution was ingested. Markers of energy metabolism, lactulose-to-rhamnose ratio, several cytokines and potential stress-related markers were measured at rest and during exercise. In addition, untargeted urine metabolite profiles were obtained. The complete procedure (Test) was repeated 1week later (Retest) to assess repeatability. Metabolic effect parameters with regard to energy metabolism and urine metabolomics were similar for both the Test and Retest period, underlining comparable exercise load. Following exercise, intestinal permeability (1h plasma lactulose-to-rhamnose ratio) and the serum interleukin-6, interleukin-10, fibroblast growth factor-21 and muscle creatine kinase concentrations were significantly increased compared with rest only during the first test and not when the test was repeated. Responses to strenuous exercise in well-trained young men, as indicated by intestinal markers and myokines, show adaptation in Test-Retest outcome. This might be attributable to a carry-over effect of the defense mechanisms triggered during the Test. This finding has implications for the design of studies aimed at evaluating physiological responses to exercise.
Background: Identification of detected compounds in untargeted LC/MS profiling is a common bottleneck in metabolomics. The CASMI contest challenges mass spectrometry experts and algorithm developers to evaluate how reliable their methods derive molecular formulae and structures from blinded mass spectral data.Objective: The application of the MAGMa software to solve the CASMI 2014 challenges is described.Methods: MAGMa was used to automatically retrieve candidate molecular structures from the HMDB and PubChem chemical databases, based on MS 1 precursor m/z values, and to provide a score indicating how well they explain the accurate MS 2 spectra.Results: For 40 out of 48 challenges, candidates with the correct molecular formula were ranked on top. For 22 out of 42 challenges the top-ranked candidate also represented the correct chemical structure and in 9 other cases the correct molecule was ranked in the top 10.Conclusion: Advantages and limitations of the approach and consequences with respect to retrieving and scoring of the correct candidates are discussed.