The development and discovery of new materials can be significantly enhanced through the adoption of FAIR (Findable, Accessible, Interoperable, and Reusable) data principles and the establishment of a robust data infrastructure in support of materials informatics. A FAIR data infrastructure and associated best practices empower materials scientists to access and make the most of a wealth of information on materials properties, structures, and behaviors, allowing them to collaborate effectively, and enable data-driven approaches to material discovery. To make data findable, accessible, interoperable, and reusable to materials scientists, we developed and are in the process of expanding a materials data infrastructure to capture, store, and link data to enable a variety of analytics and visualizations. Our infrastructure follows three key architectural design philosophies: (i) capture data across a federated storage layer to minimize the storage footprint and maximize the query performance for each data type, (ii) use a knowledge graph-based data fusion layer to provide a single logical interface above the federated data repositories, and (iii) provide an ensemble of FAIR data access and reuse services atop the knowledge graph to make it easy for materials scientists and other domain experts to explore, use, and derive value from the data. This paper details our architectural approach, open-source technologies used to build the capabilities and services, and describes two applications through which we have successfully demonstrated its use. In the first use case, we created a system to enable additive manufacturing data storage and process parameter optimization with a range of user-friendly visualizations. In the second use case, we created a system for exploring data from cathodic arc deposition experiments to develop a new steam turbine coating material, fusing a combination of materials data with physics-based equations to enable advanced reasoning over the combined knowledge using a natural language chatbot-like user interface.
Structured data artifacts such as tables are widely used in scientific literature to organize and concisely communicate important statistical information. Discovering relevant information in these tables remains a significant challenge owing to their structural heterogeneity, dense and often implicit semantics, and diffuse context. This paper describes how we leverage semantic technologies to enable technical experts to search and explore tabular data embedded within scientific documents. We present a system for the on-demand construction of knowledge graphs representing scientific tables (drawn from online scholarly articles hosted by PubMed Central), and for synthesizing tabular responses to semantic search requests against such graphs. We discuss key differentiators in our overall approach, including a two-stage semantic table interpretation that relies on an extensive structural and syntactic characterization of scientific tables, and a prototype knowledge discovery engine that uses automatically-inferred semantics of scientific tables to serve search requests by potentially fusing information from multiple tables on the fly. We evaluate our system on a real-world dataset of approximately 120,000 tables extracted from over 62,000 COVID-19-related scientific articles.
Entity linking is an important step towards constructing knowledge graphs that facilitate advanced question answering over scientific documents, including the retrieval of relevant information included in tables within these documents. This paper introduces a general-purpose system for linking entities to items in the Wikidata knowledge base. It describes how we adapt this system for linking domain-specific entities, especially for those entities embedded within tables drawn from COVID-19-related scientific literature. We describe the setup of an efficient offline instance of the system that enables our entity-linking approach to be more feasible in practice. As part of a broader approach to infer the semantic meaning of scientific tables, we leverage the structural and semantic characteristics of the tables to improve overall entity linking performance.
Materials scientists are facing increasingly challenging multi-objective performance requirements to meet the needs of modern systems such as lighter-weight and more fuel-efficient aircraft engines, and higher heat and oxidation-resistant steam turbines. While so-called second wave statistical machine learning techniques are beginning to accelerate the materials development cycle, most materials science applications are data-deprived when compared to the vastness and complexity of the search space of possible solutions. In line with DARPA’s vision of third wave AI approaches, we believe a combination of data-driven statistical machine learning and domain knowledge will be required to achieve a true revolution in materials discovery. To that end, we envision and have begun reducing to practice a system that fuses three forms of knowledge—factual scientific knowledge, physics-based and/or data-driven analytical models, and domain expert knowledge—into a single ‘Compound Knowledge Graph’ in which contextual reasoning and adaptation can be performed to answer increasingly complex questions. We believe this Compound Knowledge Graph-based system can be the nucleus of a collaborative AI assistant that supports stateful natural language back-and-forth dialogs between materials scientists and the AI to accelerate the development and discovery of new materials. This paper details our vision, summarizes our progress to date on a steam turbine blade coating use case, and outlines our thoughts on the key challenges in making this vision a reality.
Scientific models hold the key to better understanding and predicting the behavior of complex systems. The most comprehensive manifestation of a scientific model, including crucial assumptions and parameters that underpin its usability, is usually embedded in associated source code and documentation, which may employ a variety of (potentially outdated) programming practices and languages. Domain experts cannot gain a complete understanding of the implementation of a scientific model if they are not familiar with the code. Furthermore, rapid research and development iterations make it challenging to keep up with constantly evolving scientific model codebases. To address these challenges, we develop a system for the automated creation and human-assisted curation of a knowledge graph of computable scientific models that analyzes a model's code in the context of any associated inline comments and external documentation. Our system uses knowledge-driven as well as data-driven approaches to identify and extract relevant concepts from code and equations from textual documents to semantically annotate models using domain terminology. These models are converted into executable Python functions and then can further be composed into complex workflows to answer different forms of domain-driven questions. We present experimental results obtained using a dataset of code and associated text derived from NASA's Hypersonic Aerodynamics website.
This manuscript presents a data quality analysis and holistic ‘machine learning-readiness’ evaluation of a representative set of large-scale, real-world phasor measurement unit (PMU) datasets provided under the United States Department of Energy-funded FOA 1861 research program. A major focus of this study is to understand the present-day suitability of large-scale, real-world synchrophasor datasets for application of commercially-available, off-the-shelf big data and supervised or semi-supervised machine learning (ML) tools and catalogue any major obstacles to their application. To this end, dataset quality is methodically examined through an interconnect-wide quantifications of basic bad data occurrences, a summary of several harder-to-detect data quality issues that can jeopardize successful application of machine learning, and an evaluation of the adequacy of event log labeling for supervised training of models used for online event classification. A global ‘six-point’ statistical analyses of several key dataset variables is demonstrated as a means by which to identify additional hard-to-detect data quality issues, also providing an example successful application of big data technology to extract insights regarding reasonable operational bounds of the US power system. Obstacles for application of commercial ML technologies are summarized, with a particular focus on supervised and semi-supervised ML. Lessons-learned are provided regarding challenges associated with present-day event labeling practices, large spatial scope of the dataset, and dataset anonymization. Finally, insight into efficacy of employed mitigation strategies are discussed, and recommendations for future work are made.
The discovery of `event signatures' and useful insights from very large historical Phasor Measurement Unit (PMU) datasets is predicated on offline Big Data analysis approaches that rely on the generation of predictive features on a massive scale. This paper presents lessons learned from a data platform perspective towards reducing barriers to adoption of Big Data analytics against a real dataset of almost half a trillion data points drawn from over 400 PMUs distributed across the North American power grid. We demonstrate software abstractions and targeted performance optimizations that can lead to significant productivity gains for power systems researchers seeking to perform offline exploratory temporal analysis and modeling tasks, with a focus on feature generation. We describe how our optimized approach goes beyond a naive application of mainstream Big Data technologies, enabling feature generation tasks, that previously took days or even weeks, to now be completed in just a few hours.
A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.
The evolution of the energy and utilities industry, comprising smart power grids and myriad distributed energy resources, bears parallels to Industry 4.0 digitalization. Replete with diverse data modalities, the grid industry has traditionally been steeped in standards defined to primarily govern information interchange between data systems and applications. We explore the practical application of Knowledge Graph technology to make these standards more actionable. This work-in-progress paper proposes and demonstrates the idea of on-demand knowledge graphs to provision clean, integrated power grid data, in accordance with IEC CIM standards, for an enterprise ML-based analytic solution targeting better response to storm-induced power outages.
The "digital twin" is emerging as a dominant paradigm aimed at improving outcomes associated with physical counterparts within several industrial sectors. Recently, the paradigm has been gaining traction in design and engineering scenarios. Critical to the creation and maintenance of a digital twin is a "digital thread" that allows the continuous collection and linking of relevant data and analytical models throughout the lifecycle of a physical entity. This paper describes GE Research's digital thread approach, comprising a federated multimodal data platform integrating modern semantic and Big Data technologies. We demonstrate the value of our digital thread platform in enabling the digital twin of an additive manufacturing process and system to accelerate time-to-market of game-changing new parts for the aviation, power and healthcare industries.
To deliver business value, most data-driven enterprises and applications require data to be extracted and merged from otherwise siloed data storage platforms. FDC Cache has been designed and developed to enable the fusion and caching of data drawn from multiple small and/or Big Data stores. This capability executes a sequence of queries, wherein the results from one query may be used to constrain subsequent queries. The results of each query are linked with results from previous queries, incrementally building a cache of semantically linked data that can be used to support multiple independent data requests. FDC Cache uses Semantic Web technologies, and knowledge graphs in particular, to describe the relevant data and relationships in a computable model. This enables applications to reason over the graph, for example to dynamically retrieve targeted subsets of data comprised of previously disparate information. We have successfully applied FDC Cache to two distinct industrial use cases: (i) merging data across multiple sources to assemble information about current parts in a gas turbine, and (ii) dynamically aligning siloed data from electric grid transmission and distribution networks to an industry-standard common model, in which the cache creation time has been shown to scale sub-linearly with the number of data elements. FDC Cache has been open-sourced as part of the GE-developed open source Semantics Toolkit.
Organizations such as GE are heavily invested in applying advanced data-intensive machine learning (ML) techniques towards continually improving the performance of complex physical assets and industrial operations. This paper highlights some unique industrial data management challenges and demonstrates the need for and benefits of a knowledge-driven approach to data management that complements existing efforts by the ML systems community. Specifically, we present a novel software abstraction called a NodeGroup for accessing (within some domain-specific context) heterogeneous data so that the development of end-to-end ML-driven applications is further streamlined. We present our preliminary use of NodeGroups for ML applications within a prototype (additive) manufacturing data platform at GE.
A novel framework including experimental and model-based techniques saves time and enables the introduction of new alloys for additive manufacturing. This article describes the first phase of the probabilistic machine learning framework that was successfully demonstrated to rapidly define optimum parameter sets for commercial high-temperature nickel superalloys, as well as to guide alloy design and selection for compatibility with laser powder bed fusion additive manufacturing.
Additive technologies are expected to revolutionize manufacturing across almost every industry, but there is a sizable gap between the current state of the technology and the maturity required for it to achieve widespread adoption. Advanced analytics are required to improve the reliability and repeatability of additive manufacturing, and those analytics require data. Large volumes of multimodal data are generated and used throughout the additive manufacturing lifecycle, from material design to part design and simulation, part printing to post-processing and inspection. To capture and link that diverse Big Data together, we have designed and developed a federated multimodal Big Data storage and analytics platform comprised of three tiers-a distributed polyglot data storage and analysis tier with different repositories for different data structures, a metadata knowledge graph tier for modeling the data and their relationships across the various repositories, and a user interface tier for visualizing, exploring and invoking analytics on the data. The platform has been used to integrate a collection of previously disparate steps to optimize the process parameters used to build a new additive material, enabling materials scientists and other non-software experts to capture, visualize and analyze the requisite data through a single user interface.
The recent successes of commercial cognitive and AI applications have cast a spotlight on knowledge graphs and the benefits of consuming structured semantic data. Today, knowledge graphs are ubiquitous to the extent that organizations often view them as a "single source of truth" for all of their data and other digital artifacts. In most organizations, however, Big Data comes in many different forms including time series, images, and unstructured text, which often are not suitable for efficient storage within a knowledge graph. This paper presents the Semantics Toolkit (SemTK), a framework that enables access to polyglotpersistent Big Data stores while giving the appearance that all data is fully captured within a knowledge graph. SemTK allows data to be stored across multiple storage platforms (e.g., Big Data stores such as Hadoop, graph databases, and semantic triple stores) - with the best-suited platform adopted for each data type - while maintaining a single logical interface and point of access, thereby giving users a knowledge-driven veneer across their data. We describe the ease of use and benefits of constructing and querying polystore knowledge graphs with SemTK via four industrial use cases at GE.
Most organizations are becoming increasingly data-driven, often processing data from many different sources to enable critical business operations. Beyond the well-addressed challenge of storing and processing large volumes of data, financial institutions in particular are increasingly subject to federal regulations requiring high levels of accountability for the accuracy and lineage of this data. For companies like GE Capital, which maintain data across a globally interconnected network of thousands of systems, it is becoming increasingly challenging to capture an accurate understanding of the data flowing between those systems. To address this problem, we designed and developed a concept lineage tool allowing organizational data flows to be modeled, visualized and interactively explored. This tool has novel features that allow a data flow network to be contextualized in terms of business-specific metadata such as the concept, business, and product for which it applies. Key analysis features have been implemented, including the ability to trace the origination of particular datasets, and to discover all systems where data is found that meets some user-defined criteria. This tool has been readily adopted by users at GE Capital and in a short time has already become a business-critical application, with over 2,200 data systems and over 1,000 data flows captured.
The era of precision medicine is best exemplified by the growing reliance on next-generation sequencing (NGS) technologies to provide improved disease diagnosis and targeted therapeutic selection. Well-established NGS data analysis software tools, in their unmodified form, can take days to identify and interpret single nucleotide and structural variations in DNA for a single patient. To improve sample analysis throughput, we developed a highly parallel end-to-end next-generation DNA sequencing data analysis pipeline in Hadoop. In our pipeline, each step is parallelized not only across samples but also within each individual sample, achieving a 30× speedup over a single server workflow execution. Furthermore, we extensively evaluate the viability of having our Hadoop-based pipeline as part of a larger commercial genomic services offering-we demonstrate how our pipeline scales sub-linearly both with the number of samples being analyzed and with the depth of coverage of those samples. In particular, on our commodity cluster, 10× as many samples resulted in only a 2.24× increase in the execution time, and a 4× increase in coverage depth resulted in only a 2.53× growth in execution time. We anticipate that such improvements will allow large cohort populations to be analyzed in parallel, and can fundamentally change the way DNA sequencing analyses are used by both researchers and clinicians.
In this paper, we show how popular open-source data analytics platforms, such as KNIME, coupled with industry standard cluster computing systems such as Hadoop and Spark can be leveraged to build a highly flexible, scalable and userfriendly collaborative framework for the analysis of high-content hyperplexed molecular imaging data in a public cloud.
Phenotypic correlations among body weights at various ages were studied using genetic stock of poultry involving White Cornish, Red Cornish and White Plymouth Rock breeds of poultry maintained at Central Poultry Farm, Patna on random mating for a large number of generations. A total of 266 birds (131 males and 135 females) were utilized to compute correlation coefficients among body weights at various ages. All the phenotypic correlations among body weights from 4 th week to 8 th week of age, both in males and females, of WC, RC and WPR purebreds were observed to be positive, highly significant (P<0.01) and of high magnitude ranging from 0.387(between 4 th week X 8 th week body weights in females of WC) to 0.976 (between 6 th week X 7 th week body weights in males of WPR) in this study. Besides, the magnitudes of standard errors of phenotypic correlations were also observed to be very low suggesting high precision of the estimates. Besides, it was found that the magnitude of phenotypic correlations of day old chick weight, in general, had a declining tendency with that of body weights at subsequent ages. This might be, possibly, due to the dilution of maternal influence as the age advances.
Sivaramakrishnan Narayanan合作论文数Biomedical Informatics, The Ohio State University "Work In Progress - Imaging"4
Ghaleb Abdulla合作论文数Lawrence Livermore National Laboratory2