Data Science and Informatics methods have been at the center of many recent scientific discoveries and have opened up new frontiers in many areas of scientific inquiry. In this talk, I will take you through some of the most recent and exciting discoveries we've made and how informatics methods planned a central role in these discoveries. First, we will look at our work on data-driven biosignature detection, specifically how we combine pyrolysis-gas chromatography-mass spectrometry and machine learning to build an agnostic molecular biosignature detection model. Next, we will talk about how we used association analysis to predict the locations of as-yet-unknown mineral deposits on Earth and potentially Mars. These advances hold the potential to unlock new avenues of economic growth and sustainable development.Finally, we will set our sights on exoplanets—celestial bodies orbiting distant stars. The discovery of thousands of exoplanets in recent years has fueled the quest to understand their formation, composition, and potential habitability. We develop informatics approaches to better understand, classify and predict the occurrence of exoplanets by embracing the complexity and multidimensionality of exoplanets and their host stars.
The locations of minerals and mineral-forming environments, despite being of great scientific importance and economic interest, are often difficult to predict due to the complex nature of natural systems. In this work, we embrace the complexity and inherent "messiness" of our planet's intertwined geological, chemical, and biological systems by employing machine learning to characterize patterns embedded in the multidimensionality of mineral occurrence and associations. These patterns are a product of, and therefore offer insight into, the Earth's dynamic evolutionary history. Mineral association analysis quantifies high-dimensional multicorrelations in mineral localities across the globe, enabling the identification of previously unknown mineral occurrences, as well as mineral assemblages and their associated paragenetic modes. In this study, we have predicted (i) the previously unknown mineral inventory of the Mars analogue site, Tecopa Basin, (ii) new locations of uranium minerals, particularly those important to understanding the oxidation-hydration history of uraninite, (iii) new deposits of critical minerals, specifically rare earth element (REE)- and Li-bearing phases, and (iv) changes in mineralization and mineral associations through deep time, including a discussion of possible biases in mineralogical data and sampling; furthermore, we have (v) tested and confirmed several of these mineral occurrence predictions in nature, thereby providing ground truth of the predictive method. Mineral association analysis is a predictive method that will enhance our understanding of mineralization and mineralizing environments on Earth, across our solar system, and through deep time.
Minerals are information-rich materials that offer researchers a glimpse into the evolution of planetary bodies. Thus, it is important to extract, analyze, and interpret this abundance of information to improve our understanding of the planetary bodies in our solar system and the role our planet's geosphere played in the origin and evolution of life. Over the past several decades, data-driven efforts in mineralogy have seen a gradual increase. The development and application of data science and analytics methods to mineralogy, while extremely promising, has also been somewhat ad hoc in nature. To systematize and synthesize the direction of these efforts, we introduce the concept of "Mineral Informatics," which is the next frontier for researchers working with mineral data. In this paper, we present our vision for Mineral Informatics and the X-Informatics underpinnings that led to its conception, as well as the needs, challenges, opportunities, and future directions of the field. The intention of this paper is not to create a new specific field or a sub-field as a separate silo, but to document the needs of researchers studying minerals in various contexts and fields of study, to demonstrate how the systemization and enhanced access to mineralogical data will increase cross- and interdisciplinary studies, and how data science and informatics methods are a key next step in integrative mineralogical studies.
Minerals are the oldest surviving materials from the formation of our solar system. They are time capsules that store and provide information about the evolution of Earth and other planetary bodies. In addition to being a cornerstone of geoscience research, minerals also have economic, industrial and commercial importance in many sectors of society. One of the fundamental questions in mineralogy and geosciences in general is “Where to find minerals?”. Due to the complex and intertwined nature of natural systems, it has been hard to predict the occurrences of minerals. However, with increase in the volume and accuracy of mineral data and rise of mineral informatics, data science and analytics methods can be developed to answer this fundamental question in mineralogy. In this contribution, we present “mineral association analysis”, a method to: 1) Predict the mineral inventory for any existing locality. 2) Predict previous unknown localities for any given mineral. Mineral association analysis is a machine learning method that uses association rule learning to find interesting patterns based on mineral occurrence data. Using mineral association analysis, we have been able to predict locations of critical minerals, such as minerals with Li- and Th-bearing phases, predict the mineral inventory of mars analogue sites, and even understand how mineralization and mineral associations changed through deep time.
Abstract Plate tectonics, as the unifying theory in Earth sciences, controls the functioning of important planetary processes on geological timescales. Here, we present an open‐source workflow that interrogates community digital plate tectonic reconstructions, primarily in the context of the planetary deep carbon cycle. We present an updated plate tectonic reconstruction covering the last 400 million years of Earth evolution and explore components of the plate–mantle system that is involved in the exchange and storage of carbon. First, the workflow enables us to estimate subduction zone lengths through time, which represent the “tap” of carbon that is released at convergent tectonic margins. Second, we explore the role of Andean‐style versus intra‐oceanic subduction regimes during Pangea assembly and breakup. Third, we provide an improved model for carbonate platform evolution since the Devonian and evaluate the interaction of subduction zones and buried carbonate platforms. Last, we present a new model for estimating oceanic age, carbon content in the upper oceanic crust, and estimated (carbon‐containing) sediment thicknesses through time and present methods to track the subduction of this material through time. These components of the deep carbon cycle are key mechanisms controlling, or at least modulating, atmospheric CO2 on geological timescales and hence strongly influencing long‐term climate. We find that the mid to Late Cretaceous greenhouse climates were likely driven by increased subduction fluxes of volatiles and increased subduction zone interactions with carbonate platforms in the Tethyan tectonic domain. Our work highlights the importance of community digital plate tectonic reconstructions as a framework for studying key systems, such as the deep carbon cycle, that influence the life‐support mechanisms on our planet.
Ecological observations and paleontological data show that communities of organisms recur in space and time. Various observations suggest that communities largely disappear in extinction events and appear during radiations. This hypothesis, however, has not been tested on a large scale due to a lack of methods for analyzing fossil data, identifying communities, and quantifying their turnover. We demonstrate an approach for quantifying turnover of communities over the Phanerozoic Eon. Using network analysis of fossil occurrence data, we provide the first estimates of appearance and disappearance rates for marine animal paleocommunities in the 100 stages of the Phanerozoic record. Our analysis of 124,605 fossil collections (representing 25,749 living and extinct marine animal genera) shows that paleo-community disappearance and appearance rates are generally highest in mass extinctions and recovery intervals, respectively, with rates three times greater than background levels. Although taxonomic change is, in general, a fair predictor of ecologic reorganization, the variance is high, and ecologic and taxonomic changes were episodically decoupled at times in the past. Extinction rate, therefore, is an imperfect proxy for ecologic change. The paleo-community turnover rates suggest that efforts to assess the ecological consequences of the present-day biodiversity crisis should focus on the selectivity of extinctions and changes in the prevalence of biological interactions.
We present the first end-to-end, transformer-based table question answering (QA) system that takes natural language questions and massive table corpora as inputs to retrieve the most relevant tables and locate the correct table cells to answer the question. Our system, CLTR, extends the current state-of-the-art QA over tables model to build an end-to-end table QA architecture. This system has successfully tackled many real-world table QA problems with a simple, unified pipeline. Our proposed system can also generate a heatmap of candidate columns and rows over complex tables and allow users to quickly identify the correct cells to answer questions. In addition, we introduce two new open domain benchmarks, E2E_WTQ and E2E_GNQ, consisting of 2,005 natural language questions over 76,242 tables. The benchmarks are designed to validate CLTR as well as accommodate future table retrieval and end-to-end table QA research and experiments. Our experiments demonstrate that our system is the current state-of-the-art model on the table retrieval task and produces promising results for end-to-end table QA.
Join open writing teams to collaborate on commentaries for a special collection describing approaches that embody synthesis, cross-disciplinary integration, and open science across geosciences.
Abstract Serpentinization refers to the alteration of ultramafic rocks that produces serpentines and secondary (hydr)oxides under hydrothermal conditions. Serpentinization can generate H2, which in turn can potentially reduce CO/CO2 and produce organic molecules via Fischer–Tropsch type (FTT) and Sabatier type reactions. Over the last two decades, serpentinization has been extensively studied in laboratories, mainly due to its potential applications in prebiotic chemistry, origin of life in extreme environments, development of carbon‐free energies and CO2 sequestration. However, the production of H2 and organics during experimental serpentinization is hugely variable from one publication to another. The experiments span over a large range of pressure and temperature conditions, and starting compositions of fluid and solid phases are also highly variable, which collectively adds up to more than a hundred variables and leads to controversial results. Therefore, it is extremely difficult to compare results between studies, explain their variability and identify key parameters controlling the reactions. To overcome these limitations, we collected and analysed 30 peer‐reviewed articles including over 100 experimental parameters and ca. 30 mineral and organic products, hence building up a database can be completed and implemented in future studies. We then extracted basic statistical information from this dataset and demonstrate how such a comprehensive dataset is essential to better interpret available data and discuss the key parameters controlling the effectiveness of H2, CH4 and other organics production during experimental serpentinization. This is essential to guide the design of future experiments.
Earth and Space Science Open Archive PosterOpen AccessYou are viewing the latest version by default [v1]Developing the Cross-Disciplinary Information Model for NASA’s Science Mission DirectorateAuthorsRuthDuerriDAhmedEleishMarkParsonsiDDanielBerriosiDKaylinBugbeePeterFoxiDSee all authors Ruth DuerriDCorresponding Author• Submitting AuthorRonin Institute for Independent ScholarshipiDhttps://orcid.org/0000-0003-4808-4736view email addressThe email was not providedcopy email addressAhmed EleishRensselaer Polytechnic Instituteview email addressThe email was not providedcopy email addressMark ParsonsiDUniversity of Alabama in HuntsvilleiDhttps://orcid.org/0000-0002-7723-0950view email addressThe email was not providedcopy email addressDaniel BerriosiDNASA Ames Research CenteriDhttps://orcid.org/0000-0003-4312-9552view email addressThe email was not providedcopy email addressKaylin BugbeeUniversity of Alabama in Huntsvilleview email addressThe email was not providedcopy email addressPeter FoxiDRensselaer Polytechnic InstituteiDhttps://orcid.org/0000-0002-1009-7163view email addressThe email was not providedcopy email address
Abstract Minerals contain important clues to understanding the complex geologic history of Earth and other planetary bodies. Therefore, geologists have been collecting mineral samples and compiling data about these samples for centuries. These data have been used to better understand the movement of continental plates, the oxidation of Earth's atmosphere and the water regime of ancient martian landscapes. Datasets found at ‘RRUFF.info/Evolution’ and ‘mindat.org’ have documented a wealth of mineral occurrences around the world. One of the main goals in geoinformatics has been to facilitate discovery by creating and merging datasets from various scientific fields and using statistical methods and visualization tools to inspire and test hypotheses applicable to modelling Earth's past environments. To help achieve this goal, we have compiled physical, chemical and geological properties of minerals and linked them to the above‐mentioned mineral occurrence datasets. As a part of the Deep Time Data Infrastructure, funded by the W.M. Keck Foundation, with significant support from the Deep Carbon Observatory (DCO) and the A.P. Sloan Foundation, GEMI (‘Global Earth Mineral Inventory’) was developed from the need of researchers to have all of the required mineral data visible in a single portal, connected by a robust, yet easy to understand schema. Our data legacy integrates these resources into a digestible format for exploration and analysis and has allowed researchers to gain valuable insights from mineralogical data. GEMI can be considered a network, with every node representing some feature of the datasets, for example, a node can represent geological parameters like colour, hardness or lustre. Exploring subnetworks gives the researcher a specific view of the data required for the task at hand. GEMI is accessible through the DCO Data Portal (https://dx.deepcarbon.net/11121/6200‐6954‐6634‐8243‐CC). We describe our efforts in compiling GEMI, the Data Policies for usage and sharing, and the evaluation metrics for this data legacy.
Large and growing data resources on the spatial and temporal diversity and distribution of the more than 400 carbon-bearing mineral species reveal patterns of mineral evolution and ecology. Recent advances in analytical and visualization techniques leverage these data and are propelling mineralogy from a largely descriptive field into one of prediction within complex, integrated, multidimensional systems. These discoveries include: (1) systematic changes in the character of carbon minerals and their networks of coexisting species through deep time; (2) improved statistical predictions of the number and types of carbon minerals that occur on Earth but are yet to be discovered and described; and (3) a range of proposed and ongoing studies related to the quantification of network structures and trends, relation of mineral “natural kinds” to their genetic environments, prediction of the location of mineral species across the globe, examination of the tectonic drivers of mineralization through deep time, quantification of preservational and sampling bias in the mineralogical record, and characterization of feedback relationships between minerals and geochemical environments with microbial populations. These aspects of Earth’s carbon mineralogy underscore the complex co-evolution of the geosphere and biosphere and highlight the possibility for scientific discovery in Earth and planetary systems.
The origin of methane and light hydrocarbons (HCs) in natural fluids from serpentinization has commonly been attributed to the abiotic reduction of oxidized carbon by H(2)through Fischer-Tropsch-type (FTT) reactions. Multiple experimental serpentinization studies attempted to identify the parameters that control the abiotic production of H-2, CH4, and light HC. H(2)is systematically and significantly formed in experiments, indicating that its production during serpentinization is well established. However, the large variance in concentration (eight orders of magnitude) is difficult to address because of the large number of parameters that vary from one experiment to another. CH(4)and light HC production is much lower and also highly variable, leading to a vivid debate on potential role of metal catalysts and organic contamination. We have built a dataset that includes experimental setups, conditions, reactants, and products from 30 peer-reviewed articles reporting on experimental serpentinization and performed dimensionality reduction and network analysis to achieve an unbiased reading of the literature and fuel the debate. Our analysis distinguishes four experimental communities that highlights usual experimental protocols and the conditions tested so far. As expected, H(2)production is mainly controlled by T and P though a strong variability remains within a given P-T range. Accessory metal-bearing phases seem to favor H(2)production, while their role as catalyst or reactant is hampered by the lack of mineralogical characterization. CH(4)and light HC concentrations are highly variable, uncorrelated to each other, and much lower than concentrations of potential reactants (H-2, initial carbon). Accessory phases proposed as FTT catalysts do not enhance CH(4)production, confirming the inefficiency of this reaction. CH(4)only displays a positive correlation with temperature suggesting a kinetic/thermal control on its forming reaction. The carbon budget of some experiments indicates contamination in agreement with available labeled(13)C studies. Salts in initial solutions are possible sources of organic contaminants. Natural systems certainly exploit longer reaction time or other reactional paths to form the observed CH(4)and HC. The reducing potential of serpentinization can also produce intermediate metastable carbon phases in liquid or solid as observed in natural samples that should be targeted in future experiments.
Earth and Space Science Open Archive PosterOpen AccessYou are viewing the latest version by default [v1]Insights from Knowledge Graphs : Introducing a new formalismAuthorsAnirudhPrabhuiDPeterFoxiDSee all authors Anirudh PrabhuiDCorresponding Author• Submitting AuthorRensselaer Polytechnic InstituteiDhttps://orcid.org/0000-0002-9921-6084view email addressThe email was not providedcopy email addressPeter FoxiDRensselaer Polytechnic InstituteiDhttps://orcid.org/0000-0002-1009-7163view email addressThe email was not providedcopy email address