This protocol describes step-by-step how to discover and annotate DNA motifs in regulatory regions of clusters/modules of co-expressed genes in plants, including non-model species. It also explains how to control the significance of the results and how to associate the discovered motifs with potentially binding transcription factors (TFs). These tasks can all be done in a web browser. While it uses the plant-dedicated mirror of the Regulatory Sequence Analysis Tools (RSAT, http://plants.rsat.eu ), it can be used on any organism hosted at any RSAT mirror.
Background Regional/national SARS-CoV-2 genomic data platforms (DP) have played a key role during the Covid-19 pandemic to centralize data, curate, process and re-share them in a consistent form and pseudonymized/anonymized to international repositories. In Europe, several countries were able to establish such infrastructures rapidly and put them in production over the course of 2021, some earlier. Methods This survey aimed to estimate the effort that was needed to establish and run these DPs during the sequencing peak of the pandemic in 2021, including activities from data curation to data brokering, and what it would take to expand these DPs to other pathogens and antimicrobial resistance from 2023 onwards. Results Overall, a median of 10 person-months (PM) were used by each DP over 2021 and a median of 18 PM (per year) would be needed to expand activities from 2023 onwards. This survey shows that short-term funding remains commonplace and a struggle for the majority of DPs. Key supporters and arguments (e.g. centered around efficiency and cost-savings) for public health authorities and research funding bodies have also been identified to help individual data platforms in strengthening their funding proposals. Conclusions Ultimately, we propose that DPs get connected into a supra-national entity to build a stronger case and get access to major infrastructure funding grants at the European and global levels.
The COVID-19 pandemic has exemplified the importance of interoperable and equitable data sharing for global surveillance and to support research. While many challenges could be overcome, at least in some countries, many hurdles within the organizational, scientific, technical and cultural realms still remain to be tackled to be prepared for future threats. We propose to (i) continue supporting global efforts that have proven to be efficient and trustworthy toward addressing challenges in pathogen molecular data sharing; (ii) establish a distributed network of Pathogen Data Platforms to (a) ensure high quality data, metadata standardization and data analysis, (b) perform data brokering on behalf of data providers both for research and surveillance, (c) foster capacity building and continuous improvements, also for pandemic preparedness; (iii) establish an International One Health Pathogens Portal, connecting pathogen data isolated from various sources (human, animal, food, environment), in a truly One Health approach and following FAIR principles. To address these challenging endeavors, we have started an ELIXIR Focus Group where we invite all interested experts to join in a concerted, expert-driven effort toward sustaining and ensuring high-quality data for global surveillance and research.
RSAT (Regulatory Sequence Analysis Tools) is a modular software suite for the analysis of cis-regulatory elements in genome sequences. Its main applications are (i) motif discovery, appropriate to genome-wide data sets like ChIP-seq, (ii) transcription factor binding motif analysis (quality assessment, comparisons and clustering), (iii) comparative genomics and (iv) analysis of regulatory variations. Nine new programs have been added to the 43 described in the 2011 NAR Web Software Issue, including a tool to extract sequences from a list of coordinates (fetch-sequences from UCSC), novel programs dedicated to the analysis of regulatory variants from GWAS or population genomics (retrieve-variationseq and variation-scan), a program to cluster mo-tifs and visualize the similarities as trees (matrix-clustering). To deal with the drastic increase of sequenced genomes, RSAT public sites have been reorganized into taxon-specific servers. The suite is well-documented with tutorials and published protocols. The software suite is available through Web sites, SOAP/WSDL Web services, virtual machines and stand-alone programs at http://www.rsat.eu/.
Unraveling the mechanisms that regulate gene expression is a complex challenge in biology. A key step to decipher the regulatory machinery is to identify cis-regulatory elements (CREs) buried in non-coding DNA sequences. CREs are short stretches of DNA that serve as binding sites for transcription factors (TFs) and are frequently summarized as sequence motifs or logos. Although much attention has been paid to model species like Arabidopsis thaliana, little is known about regulatory motifs in other ones. Here, we describe a bottom-up approach for de novo motif discovery using peach as an example. These predictions require pre-computed gene clusters grouped by their expression similarity. Proximal promoter regions were defined as four different intervals: Up 1: [-1,500 bp, +200 bp], Up 2: [-500 bp, +200 bp], Up 3: [-500 bp, 0 bp] and Up 4 [0 bp, +200 bp]. Two algorithms from RSAT::Plants (http://plants.rsat.eu) were tested (oligo and dyad analysis). Overall, 18 out of 45 co-expressed modules were enriched in motifs typical of well-known TF families (bHLH, bZip, BZR, CAMTA, DOF, E2FE, AP2-ERF, Myb-like, NAC, TCP, and WRKY) and a few uncharacterized motifs. Our results indicate that small module size and promoter window of [-500 bp, +200 bp] relative to the transcription start site (TSS) maximize the number of motifs found and reduce low-complexity signals in peach. This work yields a comprehensive collection of Prunus persica motifs without prior knowledge and provides a pipeline that can be applied to other species.
The identification of functional elements encoded in plant genomes is necessary to understand gene regulation. Although much attention has been paid to model species like Arabidopsis (Arabidopsis thaliana), little is known about regulatory motifs in other plants. Here, we describe a bottom-up approach for de novo motif discovery using peach (Prunus persica) as an example. These predictions require pre-computed gene clusters grouped by their expression similarity. After optimizing the boundaries of proximal promoter regions, two motif discovery algorithms from RSAT::Plants (http://plants.rsat.eu ) were tested (oligo and dyad analysis). Overall, 18 out of 45 co-expressed modules were enriched in motifs typical of well-known transcription factor (TF) families (bHLH, bZip, BZR, CAMTA, DOF, E2FE, AP2-ERF, Myb-like, NAC, TCP, and WRKY) and a few uncharacterized motifs. Our results indicate that small modules and promoter window of [-500 bp, + 200 bp] relative to the transcription start site (TSS) maximize the number of motifs found and reduce low-complexity signals in peach. The distribution of discovered regulatory sites was unbalanced, as they accumulated around the TSS. This approach was benchmarked by testing two different expression-based clustering algorithms (network-based and hierarchical) and, as control, genes grouped for harboring ChIPseq peaks of the same Arabidopsis TF. The method was also verified on maize (Zea mays), a species with a large genome. In summary, this article presents a glimpse of the peach regulatory components at genome scale and provides a general protocol that can be applied to other species. A Docker software container is released to facilitate the reproduction of these analyses.
Based on our analysis, and as confirmed by the global study convened by the World Health Organization (WHO) and Chinese authorities, there is as yet no evidence demonstrating a fully natural origin of this virus. The zoonosis hypothesis, largely based on patterns of previous zoonosis events, is only one of a number of possible SARS-CoV-2 origins, alongside the research-related accident hypothesis.
Studies have shown that severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) can appear in feces within three days of infection, which is much sooner than the time taken for people to develop symptoms and get an official diagnosis (Mallapaty 2020). Therefore, earlier identification in wastewater of the virus’s arrival in a community might limit the health and economic damage caused by the coronavirus disease 2019 (COVID-19). Despite centuries of technological advances, the COVID-19 pandemic is highlighting the limitations of our ability to control viral spread and infections worldwide. This knowledge gap has induced highly variable and sometimes contradictory government decisions in many countries, such as whether to wear a mask or not, and various strategies for testing, contact tracing and isolation of infected patients. The failure of control strategies has led many countries to lockdowns. In particular, powerful and rapid analytical techniques are actually still missing because virus detection is a key step in controlling viral transmission. The biomarker approach, which was initially developed in geochemistry to measure the maturity of dead organic matter and to link molecular fossils with biological homologs, has been applied to modern samples such as soils, sediments, wastewaters and plants (Albrecht and Ourisson 1971; Lichtfouse et al. 1994 1997; Bryselbout et al. 1998; Payet et al. 1999; Dsikowitzky and Schwarzbauer 2014; Choi et al. 2020). Here, we discuss the emergence of wastewater epidemiology based on ribonucleic acid (RNA) biomarkers, a discipline linking biomedical and environmental sciences for the detection of pathogens, with focus on analytical challenges.
BACKGROUND:Life scientists routinely face massive and heterogeneous data analysis tasks and must find and access the most suitable databases or software in a jungle of web-accessible resources. The diversity of information used to describe life-scientific digital resources presents an obstacle to their utilization. Although several standardization efforts are emerging, no information schema has been sufficiently detailed to enable uniform semantic and syntactic description-and cataloguing-of bioinformatics resources. FINDINGS:Here we describe biotoolsSchema, a formalized information model that balances the needs of conciseness for rapid adoption against the provision of rich technical information and scientific context. biotoolsSchema results from a series of community-driven workshops and is deployed in the bio.tools registry, providing the scientific community with >17,000 machine-readable and human-understandable descriptions of software and other digital life-science resources. We compare our approach to related initiatives and provide alignments to foster interoperability and reusability. CONCLUSIONS:biotoolsSchema supports the formalized, rigorous, and consistent specification of the syntax and semantics of bioinformatics resources, and enables cataloguing efforts such as bio.tools that help scientists to find, comprehend, and compare resources. The use of biotoolsSchema in bio.tools promotes the FAIRness of research software, a key element of open and reproducible developments for data-intensive sciences.
The completion of the human genome sequence triggered worldwide efforts to unravel the secrets hidden in its deceptively simple code. Numerous bioinformatics projects were undertaken to hunt for genes, predict their protein products, function and post-translational modifications, analyse protein-protein interactions, etc. Many novel analytic and predictive computer programmes fully optimised for manipulating human genome sequence data have been developed, whereas considerably less effort has been invested in exploring the many thousands of other available genomes, from unicellular organisms to plants and non-human animals. Nevertheless, a detailed understanding of these organisms can have a significant impact on human health and well-being.New advances in genome sequencing technologies, bioinformatics, automation, artificial intelligence, etc., enable us to extend the reach of genomic research to all organisms. To this aim gather, develop and implement new bioinformatics solutions (usually in the form of software) is pivotal. A helpful model, often used by the bioinformatics community, is the so-called hackathon. These are events when all stakeholders beyond their disciplines work together creatively to solve a problem. During its runtime, the consortium of the EU-funded project AllBio - Broadening the Bioinformatics Infrastructure to cellular, animal and plant science - conducted many successful hackathons with researchers from different Life Science areas. Based on this experience, in the following, the authors present a step-by-step and standardised workflow explaining how to organise a bioinformatics hackathon to develop software solutions to biological problems.
ABSTRACTGenerating meaningful interpretations of gene lists remains a challenge for all large-scale studies. Many approaches exist, often based on evaluating gene enrichment among pre-determined gene classes. Here, we conceived and implemented yet another analysis tool (YAAT), specifically for data from the widely-used model organism C. elegans. YAAT extends standard enrichment analyses, using a combination of co-expression data and profiles of phylogenetic conservation, to identify groups of functionally-related genes. It additionally allows class clustering, providing inference of functional links between groups of genes. We give examples of the utility of YAAT for uncovering unsuspected links between genes and show how the approach can be used to prioritise genes for in-depth study. Our analyses revealed several limitations to the meaningful interpretation of gene lists, specifically related to data sources and the “universe” of gene lists used. We hope that YAAT will represent a model for integrated analysis that could be useful for large-scale exploration of biological function in other species.
One year after the onset of the COVID-19 pandemic, the origin of SARS-CoV-2 still eludes humanity. Early publications firmly stated that the virus was of natural origin, and the possibility that the virus might have escaped from a lab was discarded in most subsequent publications. However, based on a re-analysis of the initial arguments, highlighted by the current knowledge about the virus, we show that the natural origin is not supported by conclusive arguments, and that a lab origin cannot be formally discarded. We call for an opening of peer-reviewed journals to a rational, evidence-based and prejudice-free evaluation of all the reasonable hypotheses about the virus' origin. We advocate that this debate should take place in the columns of renowned scientific journals, rather than being left to social media and newspapers.
We write on behalf of our coauthors1 to agree with Jacques van Helden and colleagues that scientists "need to evaluate all hypotheses on a rational basis, and to weigh their likelihood based on facts and evidence, devoid of speculation concerning possible political impacts". Scientific knowledge is essential to effectively guide future efforts to reduce the chance of another pandemic,1,2 including by mitigating or blocking all relevant pathways for a pathogen to host-shift from natural hosts to humans.
Identification of functional regulatory elements encoded in plant genomes is a fundamental need to understand gene regulation. While much attention has been given to model species as Arabidopsis thaliana , little is known about regulatory motifs in other plant genera. Here, we describe an accurate bottom-up approach using the online workbench RSAT::Plants for a versatile ab-initio motif discovery taking Prunus persica as a model. These predictions rely on the construction of a co-expression network to generate modules with similar expression trends and assess the effect of increasing upstream region length on the sensitivity of motif discovery. Applying two discovery algorithms, 18 out of 45 modules were found to be enriched in motifs typical of well-known transcription factor families (bHLH, bZip, BZR, CAMTA, DOF, E2FE, AP2-ERF, Myb-like, NAC, TCP, WRKY) and a novel motif. Our results indicate that small number of input sequences and short promoter length are preferential to minimize the amount of uninformative signals in peach. The spatial distribution of TF binding sites revealed an unbalanced distribution where motifs tend to lie around the transcriptional start site region. The reliability of this approach was also benchmarked in Arabidopsis thaliana , where it recovered the expected motifs from promoters of genes containing ChIPseq peaks. Overall, this paper presents a glimpse of the peach regulatory components at genome scale and provides a general protocol that can be applied to many other species. Additionally, a RSAT Docker container was released to facilitate similar analyses on other species or to reproduce our results. One sentence summary Motifs prediction depends on the promoter size. A proximal promoter region defined as an interval of -500 bp to +200 bp seems to be the adequate stretch to predict de novo regulatory motifs in peach
Pierre Dupont合作论文数Universit?? Catholique de Louvain5