The query optimization phase within a database management system (DBMS) ostensibly finds the fastest query execution plan from a potentially large set of enumerated plans, all of which correctly compute the same result of the specified query. Sometimes the cost-based optimizer selects a slower plan, for a variety of reasons. Previous work has focused on increasing the performance of specific components, often a single operator, within an individual DBMS. However, that does not address the fundamental question: from where does this suboptimality arise, across DBMSes generally? In particular, the contribution of each of many possible factors to DBMS suboptimality is currently unknown. To identify the root causes of DBMS suboptimality, we first introduce the notion of empirical suboptimality of a query plan chosen by the DBMS, indicated by the existence of a query plan that performs more efficiently than the chosen plan, for the same query. A crucial aspect is that this can be measured externally to the DBMS, and thus does not require access to its source code. We then propose a novel predictive model to explain the relationship between various factors in query optimization and empirical suboptimality. Our model associates suboptimality with the factors of complexity of the schema, of the underlying data on which the query is evaluated, of the query itself, and of the DBMS optimizer. The model also characterizes concomitant interactions among these factors. This model induces a number of specific hypotheses that were tested on multiple DBMSes. We performed a series of experiments that examined the plans for thousands of queries run on four popular DBMSes. We tested the model on over a million of these query executions, using correlational analysis, regression analysis, and causal analysis, specifically Structural Equation Modeling (SEM). We observed that the dependent construct of empirical suboptimality prevalence correlates positively with nine specific constructs characterizing four identified factors that explain in concert much of the variance of suboptimality of two extensive benchmarks, across these disparate DBMSes. This predictive model shows that it is the common aspects of these DBMSes that predict suboptimality, not the particulars embedded in the inordinate complexity of each of these DBMSes. This paper thus provides a new methodology to study mature query optimizers, identifies underlying DBMS-independent causes for the observed suboptimality, and quantifies the relative contribution of each of these causes to the observed suboptimality. This work thus provides a roadmap for fundamental improvements of cost-based query optimizers.
The query optimization phase within a database management system (DBMS) ostensibly finds the fastest query execution plan from a potentially large set of enumerated plans, all of which correctly compute the specified query. Occasionally the cost-based optimizer selects a slower plan, for a variety of reasons. We introduce the notion of empirical suboptimality of a query plan chosen by the DBMS, indicated by the existence of a query plan that performs more efficiently than the chosen plan, for the same query. From an engineering perspective, it is of critical importance to understand the prevalence of suboptimality and its causal factors. We examined the plans for thousands of queries run on four DBMSes, resulting in over a million query executions. We previously observed that the construct of empirical suboptimality prevalence positively correlated with the number of operators in the DBMS. An implication is that as operators are added to a DBMS, the prevalence of slower queries will grow. Through a novel experiment that examines the plans on the query/cardinality combinations, we present evidence for a previously unknown upper bound on the number of operators a DBMS may be able to support before performance suffers. We show that this upper bound may have already been reached.
We describe the Arizona-NOIRLab Temporal Analysis and Response to Events System (ANTARES), a software instrument designed to process large-scale streams of astronomical time-domain alerts. With the advent of large-format CCDs on wide-field imaging telescopes, time-domain surveys now routinely discover tens of thousands of new events each night, more than can be evaluated by astronomers alone. The ANTARES event broker will process alerts, annotating them with catalog associations and filtering them to distinguish customizable subsets of events. We describe the data model of the system, the overall architecture, annotation, implementation of filters, system outputs, provenance tracking, system performance, and the user interface.
Leveraging existing data sources to improve the declaration and management of authorship conflicts of interest.
With the widespread availability of health-oriented digital and virtual devices and software (apps), healthcare organizations are shifting their approaches to recognize how patients wish to communicate, manage their health, and share their health information. The shift in digital and virtual health is designed to improve healthcare access and quality—particularly in underserved populations, geographies, and specialties. This course will present the current and emerging digital and virtual health services, as well as the benefits and drawbacks of these technologies. This course will address various forms of telehealth, apps, portable devices, and remote monitoring strategies, as well as the role of artificial and augmented intelligence in enhancing digital and virtual experiences. After a broad review of the digital and virtual health field, this course will focus on evaluating, sustaining, and leading a digital or virtual program. Each lecture will discuss regulatory issues such as privacy, security, FDA review/approval, and when digital and virtual health services can be reimbursed. In addition to these regulatory issues, the course will instruct how to conduct a needs assessment, evaluate digital and virtual health products, implement different business models, and evaluate best practices for implementation and adoption. Preferred Prerequisite: HINF 4301.
With the avalanche of alerts to be delivered by Rubin Observatory's Legacy Survey of Space and Time and the limited resources for follow-up, we will need brokers to select intriguing alerts that warrant follow-up in a timely manner. At NSF's NOIRLab and University of Arizona, we are developing the Arizona-NOIRLab Temporal Analysis and Response to Events System (ANTARES, Saha et al. 2014, 2016, Narayan et al. 2018), to hunt for the rarest of the rare events in the time domain. In this work, we provide an overview of the ANTARES system, how we use real-time alerts from the ongoing Zwicky Transient Facility survey as a training set, and the way forwards to Rubin observatory.
With the widespread availability of health-oriented digital and virtual devices and software (apps), healthcare organizations are shifting their approaches to recognize how patients wish to communicate, manage their health, and share their health information. The shift in digital and virtual health is designed to improve healthcare access and quality—particularly in underserved populations, geographies, and specialties. This course will present the current and emerging digital and virtual health services, as well as the benefits and drawbacks of these technologies. This course will address various forms of telehealth, apps, portable devices, and remote monitoring strategies, as well as the role of artificial and augmented intelligence in enhancing digital and virtual experiences. After a broad review of the digital and virtual health field, this course will focus on evaluating, sustaining, and leading a digital or virtual program. Each lecture will discuss regulatory issues such as privacy, security, FDA review/approval, and when digital and virtual health services can be reimbursed. In addition to these regulatory issues, the course will instruct how to conduct a needs assessment, evaluate digital and virtual health products, implement different business models, and evaluate best practices for implementation and adoption. Preferred Prerequisite: HINF 4301.
We describe the scientific goals and capabilities of the ANTARES project. This is a software infrastructure system designed to process time-domain alerts at the scale the Large Synoptic Survey Telescope will produce. This system will allow access to large-scale time-domain streams with real-time filters, machine learning, and other tools.
Abstract: The increasing tendency across scientific disciplines to write multi authored papers [1,2] makes the issue of the sequence of contributors’ names a major topic both in terms of reflecting actual contributions and in a posteriori assessments by evaluation committees. The reviewers aware that there are different cultures to authorship order. The usual and informal practice of giving the whole credit (impact factor) to each author of a multi authored paper is not adequate and over emphasizes the minor contributions of many authors. Similarly, evaluation of authors according to citation frequencies means often overrating resulting from high-impact but multi authored publications. Teja Tscharntke et al. [72] proposed that four methods. Like as SDC,EC, FLAE, and PCI. Comparison of the credit for contributions to this study under the four different models has been suggested. The proposed systems, such as Individual Frequency (IF) and Weighted Frequency (WF), have no repeated impact for each position.
The unprecedented volume and rate of transient events that will be discovered by the Large Synoptic Survey Telescope (LSST) demand that the astronomical community update its follow-up paradigm. Alert-brokers-automated software system to sift through, characterize, annotate, and prioritize events for follow-up-will be critical tools for managing alert streams in the LSST era. The Arizona-NOAO Temporal Analysis and Response to Events System (ANTARES) is one such broker. In this work, we develop a machine learning pipeline to characterize and classify variable and transient sources only using the available multiband optical photometry. We describe three illustrative stages of the pipeline, serving the three goals of early, intermediate, and retrospective classification of alerts. The first takes the form of variable versus transient categorization, the second a multiclass typing of the combined variable and transient data set, and the third a purity-driven subtyping of a transient class. Although several similar algorithms have proven themselves in simulations, we validate their performance on real observations for the first time. We quantitatively evaluate our pipeline on sparse, unevenly sampled, heteroskedastic data from various existing observational campaigns, and demonstrate very competitive classification performance. We describe our progress toward adapting the pipeline developed in this work into a real-time broker working on live alert streams from time-domain surveys.
Philippe Bonnet合作论文数IT University of Copenhagen60
Sameh Elnikety合作论文数Microsoft Research in Cambridge49
Gerhard Weikum合作论文数Department of Databases and Information Systems, Max-Planck Institute for Informatics48
Carson Kai-Sang Leung合作论文数Database & Data Mining Lab, Department of Computer Science, University of Manitoba42