Objective Cohort selection is ubiquitous and essential, but manual and ad hoc approaches are time-consuming, labor-intense, and difficult to scale. We sought to automate the task of cohort selection by building self-service tools that enable researchers to independently generate datasets for population sciences research. Materials and Methods The California Teachers Study (CTS) is a prospective observational study of 133,477 women who have been followed continuously since 1995. The CTS includes extensive survey-based and real-world data from cancer, hospitalization, and mortality linkages. We curated data from our data warehouse into a column-oriented database and developed a researcher-facing web application that guides researchers through the project lifecycle; captures researchers’ inputs; and automatically generates custom and analysis-ready data, code, dictionaries, and documentation. Results Researchers can register, access data, and propose projects on the CTS Researcher Platform via our CTS website. The Platform supports cohort and cross-sectional study designs for cancer, mortality, and any other ICD-based phenotypes or endpoints. User-friendly prompts and menus capture analytic design, inclusion/exclusion criteria, endpoint definitions, censoring rules, and covariate selection. Our platform empowers researchers everywhere to query, choose, review, and automatically and quickly receive custom data, analytic scripts, and documentation for their research projects. Research teams can review, revise, and update their choices anytime. Discussion We replaced inefficient traditional cohort-selection processes with an integrated self-service approach that simplifies and improves cohort selection for all stakeholders. Compared with manual methods, our solution is faster and more scalable, user-friendly, and collaborative. Other studies could re-configure our individual database, project-tracking, website, and data-delivery components for their own specific needs, or they could utilize other widely available solutions (e.g., alternative database or project-tracking tools) to enable similarly automated cohort-selection in their own settings. Our comprehensive and flexible framework could be adopted to improve cohort selection in other population sciences and observational research settings.
Abstract Background: Cancer Epidemiology Cohorts (CECs) amass vast amounts of participant health, lifestyle, environmental, genetic, and biologic data. Sharing these well-annotated participant data supports research across diverse cancer outcomes. CEC data often contain PHI, which limits use of open-access databases that have become a boon to other fields. Instead, most CECs create custom datasets for each project, which is labor intensive and hinders data sharing. Objective: The California Teachers Study (CTS), a prospective CEC of 133,477 women that began in 1995, sought to address this issue by building a user-friendly tool for researchers to independently and flexibly choose CTS data—i.e., to execute cohort selection—in a web-based platform. We also aimed to maintain existing data privacy and security; support a full range of study designs, exposures, and outcomes; and provide real-time query result visualizations. Methods: To support these computational demands, we chose an open-source column-oriented database management system optimized for fast analysis of large volumes of data. We curated and tagged participant self-reported data for improved searching and sorting. We built conditional queries that are modified by user selections, and included free text entry to handle outliers. To address privacy, the web-based tool displays summary data in aggregate and outputs the final dataset to a secure server. Users select their cancer endpoint definitions by SEER code, site group name, or ICD code; select analysis start and end points by participant-specific event type or hard-coded date; opt in or out of various censoring rules; and select self-reported data by questionnaire, topic, or search terms. After users submit the query, a folder is created for their project in the secure CTS remote desktop environment, which contains the dataset, a custom data dictionary, and starter scripts in SAS and in R. Results: Our first fully self-service data query module supports cancer cohort analyses and requires no CTS staff intervention to provision data. Two-thirds of recent CTS data requests have been cancer cohort analyses; if this continues, CTS could experience up to a 60% reduction in staff effort spent on creating data sets. This improvement enables instant access for researchers and improves data sharing for not only cancer cohort analyses, but across the full range of research projects. Discussion: A challenge to automating complex processes like cohort selection is focusing on edge cases. By focusing on automating the single most common request type (cancer cohort analyses), we immediately added value for our staff and users. The inclusion of data request modules for currently non-automatable study designs and data domains enables the tool to function as a single channel through which all data requests flow. Other CECs may also consider automating the most common aspects of their data request processes and sharing their results. Citation Format: Jennifer L. Benbow, Emma Spielfogel, Kai Lin, Sandeep Chandra, Paul Hughes, James V. Lacey. Self-serve data and cohort selection in the California Teachers Study: A web-based tool [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 894.
Temporal text, i.e., time-stamped text data are found abundantly in a variety of data sources like newspapers, blogs and social media posts. While today's data management systems provide facilities for searching full-text data, they do not provide any simple primitives for performing analytical operations with text. This paper proposes the temporal term histograms (TTH) as an intermediate primitive that can be used for analytical tasks. We propose an algebra, with operators and equivalence rules for TTH and present a reference implementation on a relational database system.
The goal of OpenAltimetry is to provide altimetry-specific data discovery and access with a focus on ease-of-use and quick response times. OpenAltimetry currently supports NASA's laser altimeter missions: ICESat (2003~2009) and ICESat-2 (scheduled for launch in 2018) with a powerful, web-based interactive interface that provides data discovery, processing and visualization capabilities targeting both novice and expert users across different science specializations. OpenAltimetry is modelled after NASA's highly popular EOSDIS Worldview application and uses a unique solution to allow users to efficiently interact with altimetry data. For data management, we are using a hybrid solution with highly optimized tiered storage databases and object storage. the OpenAltimetry application was built cloud-ready, with its highly scalable design and service-oriented architecture. We have developed an efficient data ingestion pipeline that streamlines processing and ingestion of altimetry datasets into the OpenAltimetry system. The waveform energy (for ICESat) and photon (for ICESat-2) data are stored in their original HDF5 format for efficiency. This will be critical for getting large volumes of altimetry data (e.g. ICESat-2) into the system rapidly for access by the wider community.
We present an approach for constructing a legal knowledge-base that is sufficiently scalable to allow for large-scale corpus-level analyses. We do this by creating a polymorphic knowledge representation that includes hybrid ontologies, semistructured representations of sentences, and unsupervised statistical extraction of topics. We apply our approach to over one million judicial decision documents from Henan, China. Our knowledge-base allows us to make corpus-level queries that enable discovery, retrieval, and legal pattern analysis that shed new light on everyday law in China.
Environmental-monitoring and observatory networks currently operating or under development at the national, regional, and global scales have the potential to provide an unprecedented understanding of our natural environment and the threats that endanger it. The breadth of these networks, as well as advances in technology (e.g., from mobile devices to in situ sensors and multidimensional satellite sensor data), will result in larger volumes of data and more complex data sets than ever before. All of these networks require robust cyberinfrastructure to support their varying mission, governance, operational, and scientific objectives. In this article, we use the Tropical Ecology Assessment and Monitoring (TEAM) Network as a fully functional environmental-monitoring network case study to highlight the key cyberinfrastructure components and services that support the network. We provide valuable lessons from our experience building the TEAM Network cyberinfrastructure and suggest future improvements that have broad applicability for other observatory and monitoring networks.
Landscape degradation, soil depletion, scarcity of water and fuel resources are common threats to ecosystem services, agricultural production and human livelihoods in the developing world. Observatory or monitoring networks which focus on the dynamics of coupled human-natural systems are challenged by scarce data, data heterogeneity across multiple domains and spatial scales, and complex models required to produce meaningful sustainability indicators. An additional challenge is to visualize these complex data and indicators in a unified, easily understandable framework. This paper presents an environmental sustainability dashboard that integrates GIS data with household and plot surveys, field data and remote sensing imagery to compute a variety of metrics of ecosystem stress. The dashboard, a web-based decision support tool, is a key cyberinfrastructure component designed to satisfy the objectives of a Tanzanian agricultural and ecosystem services monitoring pilot. Based on this experience we discuss our framework and how it can be generalized for building decision support tools, and their associated cyberinfrastructure, for multi-scale, interdisciplinary monitoring networks.
We present DIA, a Web services-based infrastructure for the Discovery, Integration, and Analysis of geoscience data, tools, and services. DIA provides a collaborative environment where scientists can share their resources (e.g., geochemical data, filtering services, etc.) by registering them through well-defined ontologies. We have developed a planetary materials ontology in OWL for this purpose. The ontology is used by different geoscientists (using Web services) to explore, extract, and integrate information from different heterogeneous data sets. The DIA system is now in its final pre-release phase. It is currently made accessible to a few geoscientists for conducting usability analyses, and it will eventually be made available to the community at large through the geoscience portal (GEON) at the San Diego Supercomputer Center.
We have developed the GEONGrid system for coordinating and managing naturally distributed computing, data, and cluster resources on the cyberinfrastructure. Recently, since the use of Grid technology is still very complex for researchers and scientists, the area of Grid Portals has made excellent progress. The Grid portal system is an emerging open Grid computing environment that promises to provide users with uniform seamless access to remote computing and data resources by providing an easy to use interface to cover over the complexity of more sophisticated Grid technologies. In this paper, we present our initial efforts in the design and implementation of service components in the GEONGrid portal. These service components may be implemented as Web services that follow the conventions of service-oriented architecture design. In this approach, service components are self-contained, have a well-defined programming interface defined in WSDL, and communicate using SOAP messaging. In building a GEONGrid portal, we also use a component-based user interface design. Portlets provide the desired component model for user interfaces in the same way as Web services. Using this approach, which allows Grid portals to be built out of reusable components, has the obvious advantages of reusability and modularity. Copyright © 2007 John Wiley & Sons, Ltd.
Scientists are confronted with significant data management problems due to the large volume and high complexity of scientific data. In particular, the latter makes data integration a difficult technical challenge. In this paper, we describe our work on semantic mediation and scientific workflows and discuss how these technologies address integration challenges in scientific data management. We first give an overview of the main data integration problems that arise from heterogeneity in the syntax, structure, and semantics of data. Starting from a traditional mediator approach, we show how semantic extensions can facilitate data integration in complex, multiple-world scenarios, where data sources cover different but related scientific domains. Such scenarios are not amenable to conventional schema integration approaches. The core idea of semantic mediation is to augment database mediators and query evaluation algorithms with appropriate knowledge representation techniques to exploit information from shared ontologies. Semantic mediation relies on semantic data registration, which associates existing data with semantic information from an ontology.The KEPLER scientific workflow system addresses the problem of synthesizing, from existing tools and applications, reusable workflow components and analytical pipelines to automate scientific analyses. After presenting core features and example workflows in KEPLER, we present a framework for adding semantic information to scientific workflows. The resulting system is aware of semantically plausible connections between workflow components as well as between data sources and workflow components. This information can be used by the scientist during workflow design, and by the workflow engineer, for creating data transformation steps between semantically compatible but structurally incompatible analytical steps.
Geoscience studies produce data from various observations, experiments, and simulations at an enormous rate. With proliferation of applications and data formats, the geoscience research community faces many challenges in effectively managing and sharing resources and in efficiently integrating and analyzing the data. In this paper, we discuss how this challenge is being addressed by the GEON Portal, a Web based distributed resource management system that provides integrated access to data and tools needed for knowledge discovery in the geosciences. Unlike previous data management efforts that were either data-driven or application-driven, the GEON Portal provides facilities for efficient sharing, discovery and integration of both data and services that use geoscience data. We identify the challenges involved in managing geoscientific resources and provide solutions that exploit the syntactic, semantic, temporal and spatial metadata associated with the resources. One of our goals is to provide some insight into the challenges involved in providing a comprehensive scientific data management solution based on our experiences with geoscientific data.
Scientists are confronted with significant data management problems due to the large volume and high complexity of scientific data. In particular, the latter makes data integration a difficult technical challenge. In this paper, we describe our work on semantic mediation and scientific workflows and discuss how these technologies address integration challenges in scientific data management. We first give an overview of the main data integration problems that arise from heterogeneity in the syntax, structure, and semantics of data. Starting from a traditional mediator approach, we show how semantic extensions can facilitate data integration in complex, multiple-world scenarios, where data sources...
When trying to combine different geologic maps, a number of interoperability challenges need to be overcome. We fi rst provide an overview of those challenges and then briefl y describe a mediator architecture devised to overcome them. We then focus on the problem of providing integrated access to a set of geologic maps from different state geological surveys, by defi ning a global view on the different local source schemas. Next we address the problem of content heterogeneity by defi ning an “integration ontology” to which the various local data content are mapped. The integration ontology consists of various sub-ontologies, such as one for geologic age (Poling, 1997), and several others relating to rock classifi cation (chemical composition, texture, fabric, and genesis). The latter are derived from a recent proposal for rock classifi cation (Struik and others, 2002). Based on the integration ontology, the prototype allows the user novel “concept-based” access and querying capabilities across the different geologic maps. This system is being embedded in the service-oriented data grid infrastructure under development in GEON.
In the GEON projectwe aredevelopingan interoperabilitysystemon top of ArcIMS for registering spatialdatasetsto ontologiesandsubsequentlyqueryingregistereddatasetsthroughthe ontologiesfor maprendering.Thesystemconsistsanontologyrepository, a datasetregistrationprocedure,anda query rewriting system.User-definedontologiesareimportedasOWL filesandsavedin theontologyrepository. Structuralandsemanticheterogeneitiesof datasourcesareresolved using informationfrom the dataset registrationprocedureandontologies.Thoseareusedwhenrewriting userqueries(e.g.,a geologicageor rocktypewill expandto theircorresponding“sub-concepts”in therewrittenquery).Multiple ontologiesare supportedin thesystemby allowing usersto defineanarticulationbetweentwo ontologieswhich equates someconceptsin the sourceontologyto someconceptsin the target ontology. Usersareableto switch betweenontologiesfor whichanarticulationexists.
In many data-centric scientific applications it is common to register datasets and computational services with a federation registry (also commonly called a catalog, directory, or repository). For example, the scientific data-handling system under development in the SEEK project must consider various dataset registries, including: MCAT, for access to SRB-registered datasets Metacat, for KNB-registered datasets DiGIR, for UDDI-registered data and Xanthoria, an XML-based data registry. A challenge for SEEK, and similar efforts such as GEON is to provide uniform access to registries and registered resources, based on emerging Web and grid standards.
Following a brief introduction to classical and behavioral algebraic specification, this paper discusses the verification of behavioral properties using BOBJ, especially its implementation of conditional circular coinductive rewriting with case analysis. This formal method is then applied to proving correctness of the alternating bit protocol, in one of its less trivial versions. We have tried to minimize mathematics in the exposition, in part by giving concrete illustrations using the BOBJ system.
Experience suggests that fully automated schema matching is infeasible, especially when n-to-m matches involving semantic functions are required. It is therefore advisable for a matching algorithm not only to do as much as possible automatically, but also to accurately identify positions where user input is necessary. Our matching algorithm combines several existing approaches with a new emphasis on using the context provided by the way elements are embedded in paths. A prototype tested on biological data (gene sequence, DNA, RNA, etc.) and on bibliographic data shows significant performance improvements, and will soon be included in our DDXMI data integration system. Some ideas for further improving the algorithm are also discussed.
This paper describes some formal tools to support distributed cooperative software engineering. Workers at diierent sites can collaborate on tasks including speci-cation, reenement, proving and documentation. A design record database supports alternative and incomplete development activities, and is read using any web browser; remote proof execution, animation, and informal explanation are supported, and results are broadcast by a protocol that resolves inconsistencies while updating local databases. The Kumo tool generates project websites and assists with formal veriications. Some experiments with these tools are reported.