Modern natural products (NPs) research relies on untargeted liquid chromatography coupled with mass spectrometry metabolomics. Together with cutting-edge processing and computational annotation strategies, such approaches can yield extensive spectral and structural information. However, current processing workflows require feature-alignment steps based on retention time which hinders the comparison of samples originating from different batches or analyzed using different instrumental setups. In addition, there is currently no analytical framework available to efficiently match processed metabolomics data and associated metadata with external resources. To address these limitations, we present a new sample-centric and knowledge-driven framework allowing multi-modal data alignment - e.g. through chemical structures, biological activities, or spectral features - and demonstrate its value in exploring large and chemodiverse natural extract datasets. Here, the experimental data is processed at the sample level, matched with external identifiers when possible, semantically enriched, and integrated into a unified knowledge graph. The use of semantic web technology enables comparison of processed and standardized data, information, and knowledge at the repository scale. We demonstrate the utility of the developed framework, the Experimental Natural Products Knowledge Graph (ENPKG), to leverage the results obtained from screening 1,600 plant extracts against trypanosomatids and streamline the identification of new antiparasitic compounds. Thanks to its versatility, the proposed approach allows for a radically novel exploitation of metabolomics data. Semantic web technologies are a fundamental asset and we anticipate that their adoption will strongly complement the current computational metabolomics pipelines and enable the community to advance in the description of global chemodiversity and drug discovery projects.
Natural products (NP) have proven to be a rich source of potentially bioactive compounds, and metabolomics is the current method of choice for characterizing natural extracts. To integrate the vast amount of data and information produced by modern metabolomics workflows, we recently developed a sample-centric approach for the semantic enrichment and alignment of metabolomics datasets. The resulting Experimental Natural Products Knowledge Graph (ENPKG) is queryable and integrates both newly acquired digitalized experimental data and information, and previously reported knowledge. It allows the highlighting of putative bioactive compounds at the extract level by comparing, for example, the occurrence of compounds of a given chemical class with bioactivity results. Using this approach, we recently described potent anti-Trypanosoma cruzi activity of two rotenoids, deguelin and rotenone. These compounds were identified in six active extracts from four plant species: Cnestis palala (Connaraceae), Chadsia grevei, Pachyrhizus erosus, and Desmodium heterophylum (Fabaceae). In this work, we present the results of the phytochemical investigation of four of these extracts and the establishment of a library of structural analogs for in vitro bioactivity testing. This work led to the isolation, characterization, and biological evaluation of the anti-T. cruzi potential of 41 compounds, including 11 rotenoids and seven compounds reported for the first time. Thanks to modern metabolite annotation and single-step isolation procedures, this work also demonstrates the possibility of considering natural extract libraries as a reservoir of rapidly accessible pure NPs. This perspective could increase the options for NP research and help accelerate NP drug discovery efforts.
The rising threat of Multidrug-Resistant Tuberculosis (MDR-TB), caused by Mycobacterium tuberculosis (Mtb), underscores the urgent need for new therapeutic solutions to tackle the challenge of antibiotics resistance. The current study utilized an innovative 3R infection model featuring the amoeba Dictyostelium discoideum infected with Mycobacterium marinum, serving as stand-ins for macrophages and Mtb, respectively. This high-throughput phenotypic assay allowed for the evaluation of more specific anti-infective activities that may be less prone to resistance mechanisms. To discover novel anti-infective compounds, a diverse collection of 1,600 plant extracts from the Pierre Fabre Library (PFL) was screened using the latter assay. Concurrently, these extracts underwent untargeted UHPLC-HRMS/MS analysis. The biological screening flagged the extract from Stauntonia brunoniana as one of the anti-infective hit extracts. High-resolution HPLC micro-fractionation coupled with bioactivity profiling was employed to highlight the natural products (NPs) driving this bioactivity. Stilbenes were eventually identified as the primary active compounds in the bioactive fractions. A knowledge graph (KG) was then used to leverage the heterogeneous data integrated into it to make a rational selection of stilbene-rich extracts. Using both CANOPUS chemical classes and Jaccard similarity indices (JSIs) to compare features within the metabolome of the 1600 NEs set, 14 extracts rich in stilbenes were retrieved. Among those, the roots of Gnetum edule were flagged as possessing broader chemo-diversity in their stilbene content, along with the corresponding extract also being a strict anti-infective. Eventually, a total of 11 stilbene oligomers were isolated from G. edule and fully characterized by NMR with their absolute stereochemistry established through electronic circular dichroism (ECD). Six of these compounds are new since they possess a stereochemistry which was never described in the literature to the best of our knowledge. All of them were assessed for their anti-infective activity and (-)-Gnetuhainin M was reported as having the highest anti-infective activity with an IC50 of 22.22 μM.
Global biodiversity is declining at an ever-increasing rate. Yet effective policies to mitigate or reverse these declines require ecosystem condition data that are rarely available. Morphology-based bioassessment methods are difficult to scale, limited in scope, suffer prohibitive costs, require skilled taxonomists, and can be applied inconsistently between practitioners. Environmental DNA (eDNA) metabarcoding offers a powerful, reproducible and scalable solution that can survey across the tree-of-life with relatively low cost and minimal expertise for sample collection. However, there remains a need to condense the complex, multidimensional community information into simple, interpretable metrics of ecological health for environmental management purposes. We developed a riverine taxon-independent community index (TICI) that objectively assigns indicator values to amplicon sequence variants (ASVs), and significantly improves the statistical power and utility of eDNA-based bioassessments. The TICI model training step uses the Chessman iterative learning algorithm to assign health indicator scores to a large number of ASVs that are commonly encountered across a wide geographic range. New sites can then be evaluated for ecological health by averaging the indicator value of the ASVs present at the site. We trained a TICI model on an eDNA dataset from 53 well-studied riverine monitoring sites across New Zealand, each sampled with a high level of biological replication (n = 16). Eight short-amplicon metabarcoding assays were used to generate data from a broad taxonomic range, including bacteria, microeukaryotes, fungi, plants, and animals. Site-specific TICI scores were strongly correlated with historical stream condition scores from macroinvertebrate assessments (macroinvertebrate community index or MCI; R2 = 0.82), and TICI variation between sample replicates was minimal (CV = 0.013). Taken together, this demonstrates the potential for taxon-independent eDNA analysis to provide a reliable, robust and low-cost assessment of ecological health that is accessible to environmental managers, decision makers, and the wider community.
Understanding the distribution of hundreds of thousands of plant metabolites across the plant kingdom presents a challenge. To address this, we curated publicly available LC-MS/MS data from 19,075 plant extracts and developed the plantMASST reference database encompassing 246 botanical families, 1,469 genera, and 2,793 species. This taxonomically focused database facilitates the exploration of plant-derived molecules using tandem mass spectrometry (MS/MS) spectra. This tool will aid in drug discovery, biosynthesis, (chemo)taxonomy, and the evolutionary ecology of herbivore interactions.
Plants have a complex chemo-diversity and represent a reservoir of potential new therapeutic agents. Within a Swiss research project, six scientific research groups from different disciplines are collaborating to investigate a collection of more than 17’000 unique dried plant extracts. It aims to find new bioactive molecules and their modes of action, with for example anti-infective or pro-metabolic activities. One of the main challenges of this enterprise is the management, integration and sharing of the highly heterogeneous data that are produced by the different research groups. Among these we find (i) massive high-resolution mass spectrometry data, (ii) the numerical results of innovative chemo-informatics methods, (iii) bioassay results from experimental models of tuberculosis and obesity, and (iv) organic synthetic chemistry. Additionally, requirements for data management plan and open-source science with the FAIR principles must be met. We have established an agile pipeline to capture and structure this heterogeneous data into an RDF graph. The data content's gradual expansion and evolution throughout the project presented considerable challenges, particularly in terms of data modeling. Additionally, despite many collaborators not being RDF experts, most were technically adept at producing RDF triples relevant to their contributions. We have deployed multiple instances of a triplestore and developed an in-house custom tool (i.e. KGSteward) to synchronize their content, based on a configuration file, which is centrally managed and version-controlled using Git. This strategy gave us the flexibility required to address global project challenges in common data management effectively.
Natural products exhibit interesting structural features and significant biological activities. The discovery of new bioactive molecules is a complex process that requires high-quality metabolite profiling data to properly target the isolation of compounds of interest and enable their complete structural characterization. The same metabolite profiling data can also be used to understand chemotaxonomic links between species better. This Data Descriptor details a dataset resulting from the untargeted liquid chromatography-mass spectrometry profiling of 76 natural extracts of the Celastraceae family. The spectral annotation results and related chemical and taxonomic metadata are shared, along with proposed examples of data reuse. This data can be further studied by researchers exploring the chemical diversity of natural products. This can serve as a reference sample set for deep metabolome investigation of this chemically rich plant family.
Phytochemists are aware of the contribution of plant biodiversity in providing chemical entities to the therapeutic arsenal, but the links between biodiversity and the emergence of pandemics are less described. "Healthy" biodiverse ecosystems protect us humans against emergence of infectious diseases and transmission as exemplified by the "dilution effect" developed upon the eco-epidemiological tick-borne Lyme disease. The emergence and spread of viral pandemics is not only due to the degradation of biodiversity but also compounded by anthropogenic factors such as the intensification and acceleration of trade and intercontinental transport, unprecedented urban human concentrations, climate change, industrialisation, massification and genetic uniformity in industrial breeding, consumption of bushmeat and promiscuity in certain markets with living animals. "One-Health" holistic approach to sustainable development is an excellent option to reduce the emergence of future pandemics. It tackles global environmental, human and animal health issues while respecting climate and ecological climate objectives. Surprisingly, this type of strategy implies accepting parasites and viruses and not seeing them as enemies to be eliminated. It also calls for recognition of the role of the local and indigenous communities in taking care traditionally and sustainably of more than a quarter of the world's land area. Only by thinking globally and acting locally this way, we will be able to reduce the risk of (re)-emergences of zoonoses and parasitic diseases.
In natural products (NP) research, methods for the efficient prioritization of natural extracts (NEs) are key for discovering novel bioactive NPs. In this study a biodiverse collection of 1600 NEs, previously analyzed by UHPLC-HRMS2 metabolite profiling was screened for Wnt pathway regulation. The results of the biological screening drove the selection of a subset of 30 non-toxic NEs with an inhibitory IC50 ≤ 5 μg/mL. To increase the chance of finding structurally novel bioactive NPs, Inventa, a computational tool for automated scoring of NEs based on structural novelty was used to mine the HRMS2 analysis and dereplication results. After this, 4 out of the 30 bioactive NEs were shortlisted by this approach. The most promising sample was the ethyl acetate extract of the leaves of Hymenocardia punctata (Phyllanthaceae). Further phytochemical investigations of this species resulted in the isolation of three known prenylated flavones (3, 5, 7) and ten novel bicyclo[3.3.1]non-3-ene-2,9-diones (1, 2, 4, 6, 8-13), named Hymenotamayonins. Assessment of the Wnt inhibitory activity of these compounds revealed that two prenylated flavones and three novel bicyclic compounds showed interesting activity without apparent cytotoxicity. This study highlights the potential of combining Inventa's structural novelty scores with biological screening results to effectively discover novel bioactive NPs in large NE collections.
The ENPKG framework organizes large heterogeneous metabolomics data sets as a knowledge graph, offering exciting opportunities for drug discovery and chemodiversity characterization.
The metabolome is the biochemical basis of plant form and function, but we know little about its macroecological variation across the plant kingdom. Here, we used the plant functional trait concept to interpret leaf metabolome variation among 457 tropical and 339 temperate plant species. Distilling metabolite chemistry into five metabolic functional traits reveals that plants vary on two major axes of leaf metabolic specialization-a leaf chemical defense spectrum and an expression of leaf longevity. Axes are similar for tropical and temperate species, with many trait combinations being viable. However, metabolic traits vary orthogonally to life-history strategies described by widely used functional traits. The metabolome thus expands the functional trait concept by providing additional axes of metabolic specialization for examining plant form and function.
As privileged structures, natural products often display potent biological activities. However, the discovery of novel bioactive scaffolds is often hampered by the chemical complexity of the biological matrices they are found in. Large natural extract collections are thus extremely valuable for their chemical novelty potential but also complicated to exploit in the frame of drug-discovery projects. In the end, it is the pure chemical substances that are desired for structural determination purposes and bioactivity evaluation. Researchers interested in the exploration of large and chemodiverse extract collections should thus establish strategies aiming to efficiently tackle such chemical complexity and access these structures. Establishing carefully crafted digital layers documenting the spectral and chemical complexity as well as bioactivity results of natural extracts collections can help prioritize time-consuming but mandatory isolation efforts. In this note, we report the results of our initial exploration of a collection of 1,600 plant extracts in the frame of a drug-discovery effort. After describing the taxonomic coverage of this collection, we present the results of its liquid chromatography high-resolution mass spectrometric profiling and the exploitation of these profiles using computational solutions. The resulting annotated mass spectral dataset and associated chemical and taxonomic metadata are made available to the community, and data reuse cases are proposed. We are currently continuing our exploration of this plant extract collection for drug-discovery purposes (notably looking for novel antitrypanosomatids, anti-infective and prometabolic compounds) and ecometabolomics insights. We believe that such a dataset can be exploited and reused by researchers interested in computational natural products exploration.
Larval behaviour for many of New Zealand’s diadromous freshwater fish is inadequately described. Diadromy for many amphidromous species is not obligatory however, and where conditions are suitable, freshwater larval rearing may be facilitated. Where this occurs in lakes, opportunities to document the composition and conditions supporting larval rearing exist. Boat trawling was undertaken across nine lowland lakes in the Lower Waikato over four consecutive winters with a focus on larval galaxiids. Galaxiid larvae were captured in surface water habitats in all but one lake, with banded kōkopu (Galaxias fasciatus) and giant kōkopu (Galaxias argenteus) the most common species detected. One lake, Lake Waahi, consistently resulted in the most galaxiid captures for effort expended. Analyses of larvae from this and other lakes indicated that two sizes predominated in catches and that larger, older larvae were predominantly G. fasciatus while smaller, younger larvae were predominantly G. argenteus. Stomach contents indicated that two non-native zooplankton species predominated in the diet of larvae, the Holarctic daphnia, Daphnia galeata, and the Australian calanoid copepod Boeckella symmetrica. This study provides new information regarding the timing, movement and predicted recruitment of native fish species in this river basin that has important implications for lake and river management.
Antibiotics resistance is a clear threat to the future of current tuberculosis treatments like rifampicin, prompting the need for new treatment options in this field. While plants can offer a plethora of chemical diversity in their constitutive natural products to tackle this issue, finding potentially bioactive compounds in them has not always proven to be that simple. Classical bioactivity-guided fractionation approaches are still trendy, but they bear significant shortfalls, like their time-consuming nature as well as the ever-increasing risk of isolating known bioactive compounds. In this regard, we have developed an alternative method to the latter approach that allows for natural derivatives of a known bioactive scaffold to be efficiently targeted and isolated within a large library of plant extracts. Hence our approach allows for the anticipation of bioactive structure independently of preliminary bioassays. By relying on the chemical diversity of a set of 1,600 plant extracts analyzed by HRMS/MS, we were able to isolate and characterize several minor derivatives of a previously reported bioactive aza-anthraquinone compound from Cananga brandisiana, selected within the plant set. Assessment of bioactivity on these derivatives (especially onychine, with an IC50 value of 39 µM in infection) confirmed their expected activity on Mycobacterium marinum in our anti-infective assay. This proof-of-concept study has established an original path towards bioactive compounds isolation, with the advantage of potentially highlighting minor bioactive compounds, whose activity may not even be detectable at the extract level.
This study examines the shape of scales from eleven fish species belonging to four fish families to infer whether the family, species and the geographic origin of fishes could be determined using scale shape. Site differentiation was analyzed only for the Cyprinidae since from the five species of this family three occurred in New Zealand and two in Turkey. Morphometric analysis was used because it allows standard multivariate analyses while preserving information about scale shape. Generalized Procrustes Analysis was used to analyse the data on scale shape. Principal components scores were submitted to canonical discriminant analysis to determine the efficacy of discrimination by families, species and geographic variants. The significance of classifications was assessed by MANOVA. MANOVA showed differences in the scale shape for the geographic location as well as by families and species. Families, species and geographic variants explained 91.7%, 82.4% and 95.8%, of the variation respectively. Each geographic location was correctly classified in 92.9% for Turkish and 98.4% New Zealand specimens. Fish scale shape was less effective in discriminating species from distantly related members, but better when the discrimination was among fish families, and best between fish scales for the same family but different body shapes.
Symbiotic associations require a specific communication between the host and the symbiont. In marine environments, host selection is essentially driven by chemical communication using specific metabolites also known as ecomones or semiochemicals. However, the chemical nature of these stimuli remains mostly unknown in aquatic environments. This paper explores the chemical identification of the signals allowing the pea crab Dissodactylus primitivus to recognize/find its host, the sea urchin Meoma ventricosa. We used a combination of behavioral tests aiming to monitor the chemotaxis reaction of crabs under water conditioned by different chemical cues. A classical Y-tube olfactometer was used to test pure industrial chemicals closely related to classes of molecules reported in echinoid tissues (i.e., quinonic and carotenoid pigments). A newly designed Petri-dish device was also developed to test biological fluids – mucus and perivisceral liquid- collected from the host M. ventricosa. Results show that 5,8-dihydroxy-1,4- naphtoquinone; 9–10 anthraquinones and beta-carotene allow a positive chemotaxis of the crab D. primitivus. On the other hand, the perivisceral fluid of M. ventricosa was repellent to the crabs, a behavior that may lead the crabs to select a healthy host. Based on other echinoderm-crustacean associations where the ecomones allowing host selection were identified, these results suggest that D. primitivus may benefit from several semiochemicals to locate and select its host M. ventricosa.