Academic research institutions using REDCap often face challenges aligning with U.S. FDA requirements for electronic records and signatures under 21 CFR Part 11 (Part 11). A National Center for Advancing Translational Sciences(NCATS) working group developed an implementation guide for Part 11 compliance in REDCap. Within six months after release, 259 individuals representing 164 institutions accessed the guide. Individuals who downloaded the guide reported reduced vendor reliance, improved documentation, and establishment Part 11-ready REDCap instances. This working group demonstrated how collaboration between technical and regulatory experts at many peer institutions is effective in improving regulatory compliance across the research enterprise.
Environmental exposures such as fine particulate matter, ozone, and nitrogen dioxide exhibit complex spatiotemporal variability, which is characterized by episodic spikes, multi-day persistence, and intermittent sub-threshold bursts. These transient dynamics, specifically during the episodic events, are associated with acute cardiopulmonary morbidity, suicide mortality, pregnancy and birth outcomes, and also impact non-clinical domains such as athletic performance [1],[2]. However, exposure modeling in healthcare and epidemiology relies heavily on aggregated metrics (e.g., daily averages, seasonal means), which can mask the short lived yet biologically meaningful exposure events [3],[4]. While the distributed lag models partially account for temporal structure, they often rely on predefined exposure windows. Recent deep learning-based sequence models capture complex patterns but often lack interpretability and require large labeled datasets.This dissertation proposes a motif-based framework for environmental exposure modeling using shapelets - short, discriminative subsequences in time series data, originally introduced for time-series mining [5]. Rather than summarizing exposure trajectories, we transform high-resolution monitoring data into reusable spatiotemporal exposure primitives that preserve the morphology of trajectories. In this work, shapelets are treated as computable exposure phenotypes, enabling reusable, interpretable representations of exposure dynamics across datasets. I developed a scalable environmental shapelet library using data from the US Environmental Protection Agency’s Air Quality System that was collected from 2004 to 2024 [6]. Each monitor–pollutant time series is presented as a high-quality, non-redundant subsequence at a temporal scale relevant to exposure health (7- and 30-day windows). These window lengths are based on prior evidence of acute and subacute exposure patterns. They take into account weekly and monthly temporal changes while also balancing biological relevance and computational ease. The extraction framework has seasonal balancing, variance filtering to get rid of segments that aren’t relevant, exposure-aware completeness criteria, and keeping geographic provenance metadata. A web-based retrieval framework enables interactive querying, exploration, and integration with a subsequent analytic workflow.This research enhances methodological, computational, and applied research through a scalable shapelet framework, which enables a transition from aggregate metrics to interpretable temporal primitives for environmental health modeling. By preserving fine-grained environmental dynamics, this work advances exposure analytics and supports more granular modeling of environmental determinants of health.The web-based retrieval interface has been deployed, and a standardized metadata schema has been used to preserve geographic provenance. Early linkage experiments connecting exposure shapelets with professional baseball performance windows and electronic health record-derived stillbirth conditions indicate that motif-level exposure representations encapsulate transient dynamics that remain undetected in aggregated averages, necessitating additional validation. Future work will concentrate on systematic validation across datasets and comparisons with conventional exposure representations.As a doctoral student, I want to strengthen my position in healthcare informatics and identify applications for motif-based exposure modeling. Feedback from senior researchers will help shape the remaining stages of this dissertation.
Environmental exposures such as fine particulate matter (PM2.5) exhibit substantial spatial and temporal variation. Traditional analytical methods frequently consolidate time-series data into coarse summaries (e.g., averages) that can obscure transient spikes and episodic event occurrences, such as wildfire smoke, which often contribute to acute exposure burden and align with physiologically and behaviorally meaningful responses, making them crucial for downstream health analysis. To preserve these transient dynamics, we use shapelets, which are discriminative subsequences in time series data that capture local temporal structures (e.g., sharp peaks, rapid rise and fall) that are frequently averaged in studies. We present the design and development of a reusable shapelet library of environmental time-series patterns derived from U.S. Environmental Protection Agency (EPA) Air Quality System (AQS) monitor data spanning from 2004 to 2024. We construct each monitor-pollutant time series as a collection of high-quality, non-redundant shapelets at a meaningful temporal scales (7- and 30-day windows), which are selected using exposure-aware criteria. We developed an accompanying web-based system that supports interactive exploration, retrieval, and integration with downstream analytics workflows. Overall, these contributions provide an informatics framework for representing exposure primitives as reusable and computable subsequences, which enable retrieval, comparison, and feature construction for downstream spatiotemporal analysis.
Objectives/Goals: Translational researchers spend significant amounts of time finding available datasets and other research data resources for their purposes. Objectives of this program are develop and evaluate a multipronged approach to supporting researchers with existing data resources. Methods/Study Population: We established a dedicated service with expertise in data resources to increase awareness, understanding, and utilization of existing data resources. This program assists investigators and trainees discover appropriate data resources, formulate scientific problems in computable formats, advise on state-of-the-art data analytics, data management, build collaborations, mentor data users, and develop a service pipeline for streamlined data resource project management. This is accomplished through these essential functions: (1) Discover, catalog, document, and manage metadata resources, (2) train and present data resources to the research community, (3) provide individual consultations, and (4) explore and assess novel data resources. Results/Anticipated Results: In a phased approach, the data navigation program is performing outreach to the research community and integrating with existing data efforts on campus, presenting and demonstrating existing data resources, established a consultation service, and building core competencies into long-term usage and navigation of resources across campus. Evaluating the program monthly has shown an increase in various metrics for evaluating commitment and engagement including number of requests for access to data resource, consultations, publications and presentations, co-authorship, and proposals. Unawareness and inappropriate use of data resources leads to delays in performing research and potentially unnecessary duplications of efforts. Discussion/Significance of Impact: Our data navigation program has increased use of data resources in research. Next steps are to continue evaluation and further streamline informatics approaches to data discovery, abstraction, formulation, and analysis. Harmonized data resource programs are important translational science approach to foster the next generation of research.
Background The purpose of the Ambulatory Electronic Health Record (EHR) Evaluation Tool is to provide outpatient clinics with an assessment that they can use to measure the ability of the EHR system to detect and prevent common prescriber errors. The tool consists of a medication safety test and a medication reconciliation module. Objectives The goal of this study was to perform a broad evaluation of outpatient medication-related decision support using the Ambulatory EHR Evaluation Tool. Methods We performed a cross-sectional study with 10 outpatient clinics using the Ambulatory EHR Evaluation Tool. For the medication safety test, clinics were provided test patients and associated medication test orders to enter in their EHR, where they recorded any advice or information they received. Once finished, clinics received an overall percentage score of unsafe orders detected and individual order category scores. For the medication reconciliation module, clinics were asked to electronically reconcile two medication lists, where modifications were made by adding and removing medications and changing the dosage of select medications. Results For the medication safety test, the mean overall score was 57%, with the highest score being 70%, and the lowest score being 40%. Clinics performed well in the drug allergy (100%), drug dose daily (85%), and inappropriate medication combinations (74%) order categories. Order categories with the lowest performance were drug laboratory (10%) and drug monitoring (3%). Most clinics (90%) scored a 0% in at least one order category. For the medication reconciliation module, only one clinic (10%) could reconcile medication lists electronically; however, there was no clinical decision support available that checked for drug interactions. Conclusion We evaluated a sample of ambulatory practices around their medication-related decision support and found that advanced capabilities within these systems have yet to be widely implemented. The tool was practical to use and identified substantial opportunities for improvement in outpatient medication safety.
Continual Reliability ImprovementPacifiCorp has continually and steadily improved its service reliability to customers as measured by SAIDI and SAIFI.This presentation will detail the programs and practices put in place to achieve continual improvement, as well as discuss ideas for future improvement.Jake Barker serves as director of distribution engineering and area transmission planning for PacifiCorp.He also has responsibility for customer generation engineering and power quality engineering.His responsibilities include ensuring PacifiCorp's distribution and sub-transmission grid provides adequate capacity to serve customers reliably, customers' power quality is within standards and customer generation installations comply with company interconnection policy.Previous to his current role, Barker worked in asset management developing the 10 year capital plan for major projects as well as managing customer engineering services and smart grid.
Commercial Internet of Things (IoT) sensors enable continuous data collection that benefits exposomic studies. The Exposure Health Informatics Ecosystem (EHIE) is one such sensor-based informatics platform for performing multiple simultaneous exposomic studies. It captures data from networks of sensors designed to record air quality in homes of the study’s participants and neighboring areas. In such cases where sensors are continually streaming data, it is crucial to monitor, in real time, the operational status of the network and record possible anomalies. Data collected by these sensors is only useful if it is free of errors. Therefore, maintaining the proper integrity of devices requires the capture of all deployment events that can cause anomalies. Tracking faults by recording system metadata is a difficult task, and we need a mechanism to capture the trajectories of devices within and across studies, systematically capture metadata of deployed version, and assign appropriate provenance to data recorded from each sensor. In this paper, we propose the use of a permissioned blockchain to manage the metadata and connect seemingly unrelated changes to create a trajectory of events that could result in the errors we observe. We implement a preliminary version of our blockchain solution in Hyperledger Fabric to help track errors in such a volatile setup. We also highlight how the properties of blockchain fulfill the essential needs for a metadata management solution needed in our case study.
POSTER ABSTRACTS from Third Annual Public Meeting: Mobilizing Computable Biomedical Knowledge (MCBK 2020)
Diabetes is a chronic disease with complications related to the autonomic nervous system (ANS) that can affect quality of life and lead to mortality. Clinicians and researchers currently rely on subjective and/or invasive means that don’t necessarily translate to real-world setting when assessing severity of certain diabetes complications. We elicited use-cases of studies aimed at understanding ANS in the context of diabetes to gather system requirements for designing an architecture to support sensor-based studies. Real-world studies would need to be capable of gathering contextual data as well as proxies for ANS symptoms from digital markers from an evolving sensor landscape, while also supporting the data needs of researchers before, during, and after data acquisition. The proposed architecture makes use of open source and commercially available mobile health technologies, and informatics platforms to meet the design criteria. Building and testing a prototype of the proposed architecture is planned to confirm the system performs as expected.
Exposomic research requires the generation of comprehensive spatio-temporal records of exposures along with capturing associated metadata describing limitations and uncertainties associated with the data. We describe the architecture of a metadata-driven Big Data integration platform for integration of sensor and health data to support diverse translational exposomic research.
Background and objective: In recent years, several data quality conceptual frameworks have been proposed across the Data Quality and Information Quality domains towards assessment of quality of data. These frameworks are diverse, varying from simple lists of concepts to complex ontological and taxonomical representations of data quality concepts. The goal of this study is to design, develop and implement a platform agnostic computable data quality knowledge repository for data quality assessments. Methods: We identified computable data quality concepts by performing a comprehensive literature review of articles indexed in three major bibliographic data sources. From this corpus, we extracted data quality concepts, their definitions, applicable measures, their computability and identified conceptual relationships. We used these relationships to design and develop a data quality meta-model and implemented it in a quality knowledge repository. Results: We identified three primitives for programmatically performing data quality assessments: data quality concept, its definition, its measure or rule for data quality assessment, and their associations. We modeled a computable data quality meta-data repository and extended this framework to adapt, store, retrieve and automate assessment of other existing data quality assessment models. Conclusion: We identified research gaps in data quality literature towards automating data quality assessments methods. In this process, we designed, developed and implemented a computable data quality knowledge repository for assessing quality and characterizing data in health data repositories. We leverage this knowledge repository in a service-oriented architecture to perform scalable and reproducible framework for data quality assessments in disparate biomedical data sources. (C) 2019 Elsevier B.V. All rights reserved.
Exposomic research may utilize multiple sensors to measure individuals' environment and their physiological responses. These sensors measure physical, chemical and biological properties and have wide variations in their capabilities and performance. It is therefore important to provide sensor characterization information in order to make appropriate decisions when selecting and utilizing sensors for research studies and analysis or when performing meta-studies aggregating data from multiple sensors. In this presentation, we discuss the development, organization, and use of a sensor metadata library (SML) developed by the Utah PRISMS Informatics Ecosystem (UPIE) (Grant NIH NIBIB U54EB021973).We performed a needs assessment and utilized the sensor common metadata specifications (SCMS) developed by UPIE in designing the SML. SCMS contains sensor metadata pertaining to the physical device, their deployment and resulting measurement outputs. The SML includes domains describing the physical characteristics of devices, including hardware and software versioning, measurement and/or sample collection characteristics, validation protocols, ownership and additional technical documentation. We implemented the SML using the Ne04j graph database.The SML includes tools for capturing and discovering metadata for new and updated versions of sensors. Sensor owners can submit metadata to the SML using a REDCap survey form, which is then curated and stored. Researchers can visualize stored sensor metadata graphically as interlinked nodes of information.This SML serves as a researcher-facing tool - as a repository of sensor information for researchers to design their exposomic studies and understand their limitations; and an inventory of available sensors for prospective study deployments. It also serves as a source of metadata store for the UPIE for performing semantically consistent metadata driven integration of heterogeneous sensor data streams for exposomic study analysis.
OBJECTIVES/SPECIFIC AIMS: Key factors causing irreproducibility of research include those related to inappropriate study design methodologies and statistical analysis. In modern statistical practice irreproducibility could arise due to statistical (false discoveries, p-hacking, overuse/misuse of p-values, low power, poor experimental design) and computational (data, code and software management) issues. These require understanding the processes and workflows practiced by an organization, and the development and use of metrics to quantify reproducibility. METHODS/STUDY POPULATION: Within the Foundation of Discovery – Population Health Research, Center for Clinical and Translational Science, University of Utah, we are undertaking a project to streamline the study design and statistical analysis workflows and processes. As a first step we met with key stakeholders to understand the current practices by eliciting example statistical projects, and then developed process information models for different types of statistical needs using Lucidchart. We then reviewed these with the Foundation’s leadership and the Standards Committee to come up with ideal workflows and model, and defined key measurement points (such as those around study design, analysis plan, final report, requirements for quality checks, and double coding) for assessing reproducibility. As next steps we are using our finding to embed analytical and infrastructural approaches within the statisticians’ workflows. This will include data and code dissemination platforms such as Box, Bitbucket, and GitHub, documentation platforms such as Confluence, and workflow tracking platforms such as Jira. These tools will simplify and automate the capture of communications as a statistician work through a project. Data-intensive process will use process-workflow management platforms such as Activiti, Pegasus, and Taverna. RESULTS/ANTICIPATED RESULTS: These strategies for sharing and publishing study protocols, data, code, and results across the spectrum, active collaboration with the research team, automation of key steps, along with decision support. DISCUSSION/SIGNIFICANCE OF IMPACT: This analysis of statistical methods and process and computational methods to automate them ensure quality of statistical methods and reproducibility of research.
OBJECTIVES/SPECIFIC AIMS: Issues with recruiting the targeted number of participants in a timely manner often results in underpowered studies, with more than 60% of clinical studies failing to complete or requiring extensions due to enrollment issues. The objective of this study is to develop and implement a scalable, organization wide platform to enhance accrual into clinical research studies. METHODS/STUDY POPULATION: We are developing and evaluating an informatics platform called Utah Utility for Research Recruitment (U2R2). U2R2 consists of 2 components: (i) Semantic Matcher: an automated trial criterion to patient matching component that also reports uncertainty associated with the match, and (ii) Match Delivery: mechanisms to deliver the list of matched patients for different research and clinical settings. As a first step, we limited the Semantic Matcher to utilize only structured data elements from the patient record and trial criteria. We are now including distributional semantic methods to match complete patient records and trial criteria as documents. We evaluated the first phase of U2R2 based on a randomized trial with a target enrollment of 220 participants that compares 2 treatment strategies for managing back pain (physical therapy and usual care) for individuals consulting a nonsurgical provider and symptomatic <90 days. RESULTS/ANTICIPATED RESULTS: U2R2 identified 9370 patients from the University of Utah Hospitals and Clinics as potential matches. Of these 9370, 1145 responded to the Back Pain study research team’s email or phone communications, and were further screened by phone. In total, 250 participants completed a screening visit, resulting in the current study enrollment of 130 participants. Forty-three of 1145 patients refused to participate, and 50 participants no-showed their screening visit. DISCUSSION/SIGNIFICANCE OF IMPACT: A recruitment platform can enhance potential participant identification, but requires attention to multiple issues involved with clinical research studies. Clinical eligibility criteria are usually unstructured and require human mediation and abstraction into discrete data elements for matching against patient records. In addition, key eligibility data are often embedded within text in the patient record. Distributional semantic approaches, by leveraging this content, can identify potential participants for screening with more specificity. The delivery of the list of matched patient results should consider characteristics of the research study, population, and targeted enrollment (eg, back pain being a common disorder and the possibility of the patient visiting different types of clinics), as well as organizational and socio-technical issues surrounding clinical practice and research. Embedding the delivery of match results into the clinical workflow by utilizing user-centered design approaches and involving the clinician, the clinic, and the patient in the recruitment process, could yield higher accrual indices.