Many non-natural amino acids can be incorporated by biological systems into coded functional peptides and proteins. For such incorporations to be effective, they must not only be compatible with the desired function but also evade various biochemical error-checking mechanisms. The underlying molecular mechanisms are complex, and this problem has been approached previously largely by expert perception of isomer compatibility, followed by empirical study. However, the number of amino acids that might be incorporable by the biological coding machinery may be too large to survey efficiently using such an intuitive approach. We introduce here a workflow for searching real and computed non-natural amino acid libraries for biosimilar amino acids which may be incorporable into coded proteins with minimal unintended disturbance of function. This workflow was also applied to molecules which have been previously benchmarked for their compatibility with the biological translation apparatus, as well as commercial catalogs. We report the results of scoring their contents based on fingerprint similarity via Tanimoto coefficients. These similarity scoring methods reveal candidate amino acids which could be substitutable into modern proteins. Our analysis discovers some already-implemented substitutions, but also suggests many novel ones.
During the past decade promising methods for computational prediction of electron ionization mass spectra have been developed. The most prominent ones are based on quantum chemistry (QCEIMS) and machine learning (CFM-EI, NEIMS). Here we provide a threefold comparison of these methods with respect to spectral prediction and compound identification. We found that there is no unambiguous way to determine the best of these three methods. Among other factors, we find that the choice of spectral distance functions play an important role regarding the performance for compound identification.
We report the co-polymerization of glycol nucleic acid (GNA) monomers with unsubstituted and substituted dicarboxylic acid linkers under plausible early Earth aqueous dry-down conditions. Both linear and branched co-polymers are produced. Mechanistic aspects of the reaction and potential roles of these polymers in prebiotic chemistry are discussed.
Recent findings, in vitro and in silico, are strengthening the idea of a simpler, earlier stage of genetically encoded proteins which used amino acids produced by prebiotic chemistry. These findings motivate a re-examination of prior work which has identified unusual properties of the set of twenty amino acids found within the full genetic code, while leaving it unclear whether similar patterns also characterize the subset of prebiotically plausible amino acids. We have suggested previously that this ambiguity may result from the low number of amino acids recognized by the definition of prebiotic plausibility used for the analysis. Here, we test this hypothesis using significantly updated data for organic material detected within meteorites, which contain several coded and non-coded amino acids absent from prior studies. In addition to confirming the well-established idea that "late" arriving amino acids expanded the chemistry space encoded by genetic material, we find that a prebiotically plausible subset of coded amino acids generally emulates the patterns found in the full set of 20, namely an exceptionally broad and even distribution of volumes and an exceptionally even distribution of hydrophobicities (quantified as logP) over a narrow range. However, the strength of this pattern varies depending on both the size and composition the library used to create a background (null model) for a random alphabet, and the precise definition of exactly which amino acids were present in a simpler, earlier code. Findings support the idea that a small sample size of amino acids caused previous ambiguous results, and further improvements in meteorite analysis, and/or prebiotic simulations will further clarify the nature and extent of unusual properties. We discuss the case of sulfur-containing amino acids as a specific and clear example and conclude by reviewing the potential impact of better understanding the chemical "logic" of a smaller forerunner to the standard amino acid alphabet.
A central question in origins of life research is how non-entailed chemical processes, which simply dissipate chemical energy because they can do so due to immediate reaction kinetics and thermodynamics, enabled the origin of highly-entailed ones, in which concatenated kinetically and thermodynamically favorable processes enhanced some processes over others. Some degree of molecular complexity likely had to be supplied by environmental processes to produce entailed self-replicating processes. The origin of entailment, therefore, must connect to fundamental chemistry that builds molecular complexity. We present here an open-source chemoinformatic workflow to model abiological chemistry to discover such entailment. This pipeline automates generation of chemical reaction networks and their analysis to discover novel compounds and autocatalytic processes. We demonstrate this pipeline's capabilities against a well-studied model system by vetting it against experimental data. This workflow can enable rapid identification of products of complex chemistries and their underlying synthetic relationships to help identify autocatalysis, and potentially self-organization, in such systems. The algorithms used in this study are open-source and reconfigurable by other user-developed workflows.
The Data Innovation and Science Cluster (DISC) is a core element of ESA's data quality strategy for the Aeolus mission, which was launched in August 2018. Aeolus provides for the first-time global observations of vertical profiles of horizontal wind information by using the first Doppler wind lidar in space. The Aeolus DISC is responsible for monitoring and improving the quality of the Aeolus aerosol and wind products, for the upgrade of the operational processors as well as for impact studies and support of data usage. It has been responsible for multiple significant processor upgrades which reduced the systematic error of the Aeolus observations drastically. Only due to the efforts of the Aeolus DISC team members prior to and after launch, the systematic error of the Aeolus wind products could be reduced to a global average below 1 m/s which was an important pre-requisite for making the data available to the public in May 2020 and for its use in operational weather prediction. In 2020, the reprocessing of earlier acquired Aeolus data, another important task of the Aeolus DISC, also started. In this way, also observations from June to December 2019 with significantly better quality could be made available to the public, and more data will follow this and next year. Without the thorough preparations and close collaboration between ESA and the Aeolus DISC over the past decade, many of these achievements would not have been possible.
Mass spectrometry (MS) can become a potentially useful instrument type for aerosol, droplet and fomite (ADF) contagion surveillance in pandemic outbreaks, such as the ongoing SARS-CoV-2 pandemic. However, this will require development of detection protocols and purposing of instrumentation for in situ environmental contagion surveillance. These approaches include: (1) enhancing biomarker detection by pattern recognition and machine learning; (2) the need for investigating viral degradation induced by environmental factors; (3) representing viral molecular data with multidimensional data transforms, such as van Krevelen diagrams, that can be repurposed to detect viable viruses in environmental samples; and (4) absorbing engineering attributes for developing contagion surveillance MS from those used for astrobiology and chemical, biological, radiological, nuclear (CBRN) monitoring applications. Widespread deployment of such an MS-based contagion surveillance could help identify hot zones, create containment perimeters around them and assist in preventing the endemic-to-pandemic progression of contagious diseases.
The chemical space of prebiotic chemistry is extremely large, while extant biochemistry uses only a few thousand interconnected molecules. Here we discuss how the connection between these two regimes can be investigated, and explore major outstanding questions in the origin of life.
Soon after its successful launch in August 2018, the spaceborne wind lidar ALADIN (Atmospheric LAser Doppler INstrument) on-board ESA’s Earth Explorer satellite Aeolus has demonstrated to provide atmospheric wind profiles on a global scale. Being the first ever Doppler Wind Lidar (DWL) instrument in space, ALADIN contributes to the improvement in numerical weather prediction (NWP) by measuring one component of the horizontal wind vector. The performance of the ALADIN instrument was assessed by a team from ESA, DLR, industry, and NWP centers during the first months of operation. The current knowledge about the main contributors to the random and systematic errors from the instrument will be discussed. First validation results from an airborne campaign with two wind lidars on-board the DLR Falcon aircraft will be shown.
Already within the first weeks after the launch of ESA's Earth Explorer mission Aeolus on 22 August 2018, the spaceborne wind lidar ALADIN (Atmospheric LAser Doppler INstrument) provided atmospheric backscatter measurements on 5 September and wind profiles on 12 September 2018. This swift availability of observations from ALADIN after launch is considered as a great success for ESA, space industry and algorithm and processor developer teams. These teams from scientific institutes, numerical weather prediction (NWP) centres, companies and ESA continuously improved and tested the retrieval algorithms and processors using sophisticated end-to-end simulation tools and experience gained with the airborne demonstrator for Aeolus for more than 15 years before launch. This cooperation from the pre-launch phase of Aeolus was extended within a new framework for exploitation activities of Earth Explorer missions named Data Innovation and Science Cluster (DISC) starting in January 2019. The Aeolus DISC activities range from instrument monitoring including calibration to algorithm refinement resulting in updates of the complete processor chain for all product levels every 6 months. DISC teams perform continuous monitoring of the product quality and provide regular reports in supports of external validation teams and ESA. Finally, wind product monitoring and impact experiments with NWP models are building an essential activity within the Aeolus DISC in order to achieve the objective of the Aeolus mission. In order to cover the broad range of activities, a multi-disciplinary team of experts, institutes and companies was established for the Aeolus DISC coordinated by DLR with ECMWF, KNMI, CNRS/Meteo-France, DoRIT, ABB, S&T and Serco. During the presentation the Aeolus instrument performance for wind products, the discovered causes of the systematic errors and their correction will be discussed. Main achievements in this area are related to the characterization and correction of enhanced dark signal levels for single hot pixels in June 2019, the identification of the harmonic error contribution caused by the varying telescope primary mirror temperature variation in September- October 2019, the error in the on-board computation of the satellite induced Doppler frequency shift, and finally the observed temporal drift of a constant bias caused by drifts in the internal reference path. An outlook to the implementation of these corrections for real-time and reprocessed data products will be given.
VirES (Virtual Workspace for Earth Observation Scientists) is a highly interactive data manipulation and retrieval interface for specific ESA Earth Explorer mission products. It is under evolutionary/operational maintenance by EOX since 2016 for the Swarm geomagnetic mission and starting 2018 it has been extended for ESA's Earth Explorer Aeolus wind profiling mission. The goal of this service is to provide users with intuitive and easy access to the mission’s level 1B, 2A, 2B and several auxiliary data products. In order to be able to understand, manage, visualize and analyze the new and complex data produced by the satellite a close collaboration between project partners DLR and DoRIT and EOX as industry partner has been established. VirES for Aeolus provides means for multi-dimensional visualization, interactive plotting and analysis and stands for a modern concept of extended access to Earth Observation (EO) data. It supports novel ways of data discovery, visualization, filtering, selection, analysis, snapshotting and downloading. The service was designed with a focus on the following areas of application: Quality analysis, calibration and validation, scientific exploitation, modeling and prediction. Specialized solutions have been developed to allow visualization of the complex and large datasets in an interactive and intuitive way. The tool has been further refined and improved since the Aeolus launch August 2018 in close collaboration with its users and project partners. VirES for Aeolus offers easy access to the mission data through ordinary web browsers via without the need for installing any specialized software.
The European Space Agency (ESA) wind mission, Aeolus, hosts the first space-based Doppler Wind Lidar (DWL) world-wide. The primary mission objective is to demonstrate the DWL technique for measuring wind profiles from space, intended for assimilation in Numerical Weather Prediction (NWP) models. The wind observations will also be used to advance atmospheric dynamics research and for evaluation of climate models. Mission spin-off products are profiles of cloud and aerosol optical properties. Aeolus was launched on 22 August 2018, and the Atmospheric LAser Doppler INstrument (Aladin) instrument switch-on was completed with first high energy output in wind mode on 4 September 2018 [1], [2]. The on-ground data processing facility worked excellent, allowing L2 product output in near-real-time from the start of the mission. First results from the wind profile product (L2B) assessment show that the winds are of very high quality, with random errors in the free Troposphere within (cloud/aerosol backscatter winds: 2.1 m/s) and larger (molecular backscatter winds: 4.3 m/s) than the requirements (2.5 m/s), but still allowing significant positive impact in first preliminary NWP impact experiments. The higher than expected random errors at the time of writing are amongst others due to a lower instrument out-and input photon budget than designed. The instrument calibration is working well, and some of the data processing steps are currently being refined to allow to fully correct instrument alignment related drifts and elevated detector dark currents causing biases in the first data product version. The optical properties spin-off product (L2A) is being compared e.g. to NWP model clouds, air quality model forecasts, and collocated ground-based observations. Features including optically thick and thin particle and hydrometeor layers are clearly identified and are being validated.
Compartmentalization is likely to have been essential for the emergence of life. Compartmentalization allows for the creation of unique chemical conditions that can be maintained out of equilibrium with the environment and the exclusion of parasites. Confining organic molecules also helps limit diffusion, increases concentration and can thus influence both the thermodynamics and kinetics of prebiotic reactions. Biology currently predominantly uses phospholipids to construct cell membranes. However, there are many other types of organic compounds that can form stable compartments in water, and many of these may have been abundant in the prebiotic environment. In this study we explore this alternative lipid chemical space by using structure enumeration algorithms to compute an exhaustive combinatorial library of surfactant molecules. We then predict the propensity of these compounds to self-assemble into membranes using quantitative structure-property relationship (QSPR) models on critical micelle concentration (CMC). Combined with critical packing parameter calculations, these models can allow identification of novel molecule types which can be experimentally assayed as candidates for the emergence of protocells.
Biology encodes hereditary information in DNA and RNA, which are finely tuned to their biological functions and modes of biological production. The central role of nucleic acids in biological information flow makes them key targets of pharmaceutical research. Indeed, other nucleic acid-like polymers can play similar roles to natural nucleic acids both in vivo and in vitro; yet despite remarkable advances over the last few decades, much remains unknown regarding which structures are compatible with molecular information storage. Chemical space describes the structures and properties of molecules that could exist within a given molecular formula or other classification system. Using structure generation methods, we explore nucleic acid analogues within the formula ranges BC3-7H5-15O2-4 and BC3-6H5-15N1-2O0-4, where B is a recognition element (e.g., a nucleobase). Other restrictions included two obligatory points of attachment for inclusion into a linear polymer and substructures predicting chemical stability. These sets contain 86,007 (CHO) and 75,309 (CHNO) compositionally isomeric structures, representing 706,568 CHO and 454,422 CHNO stereoisomers, that diversely and densely occupy this space. These libraries point toward there being large spaces of unexplored chemistry relevant to pharmacology and biochemistry and efforts to understand the origins of life.
Understanding complex (bio/geo)systems is a pivotal challenge in modern sciences that fuels a constant development of modern analytical technology, finding innovative solutions to resolve and analyse. In this introductory paper to the Faraday Discussion "Challenges in the analysis of complex natural systems", we aim to present concepts of complexity, and complex chemistry in systems subjected to biotic and abiotic transformations, and introduce the analytical possibilities to disentangle chemical complexity into its elementary parts (i.e. compositional and structural resolution) as a global integrated approach termed systems chemical analytics.
Life uses a common set of 20 coded amino acids (CAAs) to construct proteins. This set was likely canonicalized during early evolution; before this, smaller amino acid sets were gradually expanded as new synthetic, proofreading and coding mechanisms became biologically available. Many possible subsets of the modern CAAs or other presently uncoded amino acids could have comprised the earlier sets. We explore the hypothesis that the CAAs were selectively fixed due to their unique adaptive chemical properties, which facilitate folding, catalysis, and solubility of proteins, and gave adaptive value to organisms able to encode them. Specifically, we studied in silico hypothetical CAA sets of 3–19 amino acids comprised of 1913 structurally diverse α-amino acids, exploring the adaptive value of their combined physicochemical properties relative to those of the modern CAA set. We find that even hypothetical sets containing modern CAA members are especially adaptive; it is difficult to find sets even among a large choice of alternatives that cover the chemical property space more amply. These results suggest that each time a CAA was discovered and embedded during evolution, it provided an adaptive value unusual among many alternatives, and each selective step may have helped bootstrap the developing set to include still more CAAs.
VirES is a Virtual workspace for Earth-observation Scientists, a service provided by the European Space Agency (ESA). VirES has firstly been established for ESA’s magnetic field mission Swarm as “VirES for Swarm’‘ and has been extended to ESA’s atmospheric dynamics mission Aeolus, which was launched in August 2018. The service is developed by the Austrian IT company EOX in strong collaboration with missions’ scientists. VirES is a web-based service (https://aeolus.services) that enables scientists to discover, visualize, select and download data of Earth-observation (EO) missions through an easy to operate graphical user interface. ”VirES for Aeolus” will provide access to Aeolus L1B, L2A, L2B, L2C products and auxiliary data. The first version 1.0 passed acceptance tests in April 2018 and developments towards Version 1.2 are in progress. The service is planned to be accessible for public use as soon as the mission’s phase E1 is completed and first data products are released by ESA.
Since its inception six decades ago, astrobiology has diversified immensely to encompass several scientific questions including the origin and evolution of Terran life, the organic chemical composition of extraterrestrial objects, and the concept of habitability, among others. The detection of life beyond Earth forms the main goal of astrobiology, and a significant one for space exploration in general. This goal has galvanized and connected with other critical areas of investigation such as the analysis of meteorites and early Earth geological and biological systems, materials gathered by sample-return space missions, laboratory and computer simulations of extraterrestrial and early Earth environmental chemistry, astronomical remote sensing, and in-situ space exploration missions. Lately, scattered efforts are being undertaken towards the R&D of the novel and as-yet-space-unproven life-detection technologies capable of obtaining unambiguous evidence of extraterrestrial life, even if it is significantly different from Terran life. As the suite of space-proven payloads improves in breadth and sensitivity, this is an apt time to examine the progress and future of life-detection technologies.
The chemistry occurring in the universe generates a huge variety of organic compounds abiotically. Significant progress has been made in understanding the types and distributions of these compounds in various planetary, asteroidal, cometary, nebular, and molecular cloud environments. One of the most exciting recent discoveries was the detection of low-molecular-weight organic species native to the surface of comet 67P/Churyumov-Gerasimenko by the cometary sampling and composition experiment that was aboard the Philae lander of the European Space Agency. The identities of these species were estimated using a simple hand-fitting method. Here, we use a more rigorous statistical method to fit the same data and find that there is some variance between results obtained using the two methods. This paper offers recommendations to improve the numerical methods for fitting of mass spectra, which would lead to more confident identification of ambiguous-mass compounds. Such methods may also help maximize the gains of future sampling-driven space missions, to Europa, Titan, Mars and its moons, comets, and asteroids, by filling in gaps in current scientific knowledge regarding electron impact mass spectra of ambiguous-mass compounds.
The standard alphabet of the 20 genetically encoded amino acids is considered to have been selected during early evolution from a larger pool of alpha-amino acids based on its coverage of the chemical space. Chemical space is here defined by charge, size and hydrophobicity, leading to 6-tuples representing coverage, which is composed of range and evenness in these three physico-chemical properties. We summarize findings of previous studies on the adaptive properties of the 20 encoded amino acids and show how we extend these computational experiments to subsets of the standard alphabet.
Adalbert Kerber合作论文数Mathematics Department,
University of Bayreuth, Germany
Lehrstuhl II für Mathematik
33