The Sloan Digital Sky Survey has validated and made publicly available its First Data Release. This consists of 2099 square degrees of five-band (u g r i z) imaging data, 186,240 spectra of galaxies, quasars, stars and calibrating blank sky patches selected over 1360 square degrees of this area, and tables of measured parameters from these data. The imaging data go to a depth of r ≈ 22.6 and are photometrically and astrometrically calibrated to 2% rms and 100 milli-arcsec rms per coordinate, respectively. The spectra cover the range 3800–9200Å, with a resolution of 1800–2100. Further characteristics of the data are described, as are the data products themselves. Subject headings: Atlases—Catalogs—Surveys Enrico Fermi Institute, The University of Chicago, 5640 S. Ellis Ave., Chicago, IL 60637 Lawrence Berkeley National Laboratory, One Cyclotron Rd., Berkeley CA 94720-8160 Astronomy Centre, University of Sussex, Falmer, Brighton BN1 9QJ, United Kingdom Department of Physics, University of Michigan, 500 East University Ave., Ann Arbor, MI 48109 Institute for Astronomy Royal Observatory Blackford Hill Edinburgh EH9 3HJ Scotland Department of Physics, University of Pennsylvania, Philadelphia, PA 19104 Department of Physics, Applied Physics, and Astronomy, Rensselaer Polytechnic Institute, Troy, NY 12180 Lucent Technologies, 2701 Lucent Lane, Lisle, IL 60532 Department of Astronomy and Research Center for the Early Universe, School of Science, University of Tokyo, 7-3-1 Hongo, Bunkyo, Tokyo 113-0033, Japan Joseph Henry Laboratories, Princeton University, Princeton, NJ 08544 School of Natural Sciences, Institute for Advanced Study, Einstein Drive, Princeton, NJ 08540 Physics Department, Rochester Institute of Technology, 85 Lomb Memorial Drive, Rochester, NY 14623-5603 Department of Astronomy and Astrophysics, the Pennsylvania State University, University Park, PA 16802 University of Zagreb, Department of Physics, Bijenička cesta 32, 10000 Zagreb, Croatia Institute for Astronomy, 2680 Woodlawn Road, Honolulu, HI 96822 University of Wyoming, Dept. of Physics & Astronomy, Laramie, WY 82071 Department of Physics, Drexel University, Philadelphia, PA 19104 Max-Planck-Institut für extraterrestrische Physik, Giessenbachstrasse 1, D-85741 Garching, Germany Department of Astronomy, Ohio State University, Columbus, OH 43210
Many fields of science rely on relational database management systems to analyze, publish and share data. Since RDBMS are originally designed for, and their development directions are primarily driven by, business use cases they often lack features very important for scientific applications. Horizontal scalability is probably the most important missing feature which makes it challenging to adapt traditional relational database systems to the ever growing data sizes. Due to the limited support of array data types and metadata management, successful application of RDBMS in science usually requires the development of custom extensions. While some of these extensions are specific to the field of science, the majority of them could easily be generalized and reused in other disciplines. With the Graywulf project we intend to target several goals. We are building a generic platform that offers reusable components for efficient storage, transformation, statistical analysis and presentation of scientific data stored in Microsoft SQL Server. Graywulf also addresses the distributed computational issues arising from current RDBMS technologies. The current version supports load balancing of simple queries and parallel execution of partitioned queries over a set of mirrored databases. Uniform user access to the data is provided through a web based query interface and a data surface for software clients. Queries are formulated in a slightly modified syntax of SQL that offers a transparent view of the distributed data. The software library consists of several components that can be reused to develop complex scientific data warehouses: a system registry, administration tools to manage entire database server clusters, a sophisticated workflow execution framework, and a SQL parser library.
As some scientific projects begin to collect data well into the petabyte range, the technology used to analyze and retrieve such data must evolve appropriately. When presented with such a large volume, prior data delivery techniques, such as the time-honored system of simply copying all data to local storage prior to analysis, become impractical. This thesis studies the use of databases to host large-scale, scientific data. We describe early systems, such as SkyServer and CASJobs, and discuss an implementation of automatic provenance within them. Additionally, we discuss various observations regarding query patterns collected from such systems over time. Lastly, this thesis concludes with a discussion of TileDB, a novel distributed computing framework using independent shared-nothing databases. In addition to features common to many distributed systems, such as automatic parallelization and incremental fault tolerance, TileDB also combines several features that are fairly novel within the field, such as dynamic allocation of both data and work, long term, adaptive data curation and transparent integration with existing database deployments.
Multi-wavelength astronomical studies require cross-identification of detections of the same celestial objects in multiple catalogs based on spherical coordinates and other properties. Because of the large data volumes and spherical geometry, the symmetric N-way association of astronomical detections is a computationally intensive problem, even when sophisticated indexing schemes are used to exclude obviously false candidates. Legacy astronomical catalogs already contain detections of more than a hundred million objects while ongoing and future surveys will produce catalogs of billions of objects with multiple detections of each at different times. One time, pair-wise cross-identification of these large catalogs is not sufficient for many astronomical scenarios. Consequently, a novel system is necessary that can cross-identify multiple catalogs on-demand, efficiently and reliably. In this paper, we present our solution based on a cluster of commodity servers and ordinary relational databases. The cross-identification problems are formulated in a language based on SQL, but extended with special clauses. These special queries are partitioned spatially by coordinate ranges and compiled into a complex workflow of ordinary SQL queries. Workflows are then executed in a parallel framework using a cluster of servers hosting identical mirrors of the same data sets.
Catalog Archive Server Jobs (CasJobs) is an asynchronous query workbench service that lets users run unrestricted SQL queries against scientific catalog archives. After running queries in batch mode, users can save their results to a personal database called MyDB before downloading them, letting users manage their query workloads, results, and histories without causing network overloads.
The Panoramic Survey Telescope and Rapid Response System (Pan-STARRS) is the next generation of digital sky surveys that builds on the success of the Sloan Digital Sky Survey (SDSS) . The Pan-STARRS consortium is centered at the University of Hawai`i, Institute for Astronomy, and includes nine other institutions worldwide. The next generation system leverages SQL Server 2008, Windows Workflow Foundation and the Trident Scientific Workbench. This updated technology is needed to address the much larger data generated by Pan-STARRS and the need to make that data available to astronomers promptly. SDSS released survey data of about 4TB in size every 6 months; PS will have about 30TB of data per year, incrementally updated every week. The currently deployed PS1 telescope is one of the four Pan-STARRS telescopes. PS1 tests the system end to end, including the optics, cameras, image processing algorithms, data loading workflows, user facing databases, and science analysis. The PS1 survey over the next 3.5 years
As some scientific data resources grow well past the terabyte mark, so do the technical challenges of analyzing and retrieving such data. When presented with such a large volume, prior data delivery techniques, such as the time-honored system of simply copying all data to local storage prior to analysis, become impractical. We present CASJobs, an online workbench that focuses on enabling server-side exploratory and analytical tasks on large sets of data.
Author(s): Agarwal, Deborah A.; Goode, Monte; Gupchup, Jayant; Hunt, James; Ingen, Catharine van; Leonardson, Rebecca; Rodriguez, Matthew; Li, Nolan | Abstract: The wide variety of agencies collecting, storing, and publishing hydrologic data today makes it possible for scientists to gather a tremendous amount of data about a watershed simply by using the Internet. However, availability of the data is only the first step to its use in analysis. Typically the data sets that exist across a watershed are highly heterogeneous in data format, units, types, periods of measurement, frequency of measurements, quality, etc. In this presentation we will describe the Scientific Data Server we have designed to enable the combination of data from across a watershed into a database organized using a unifying schema. This data server has been prototyped using data from the Russian River and Bear River watersheds and is currently being used to study the hydrology and characteristics of the Russian River.
OpenSkyNode and ADQL are the major new steps in the Data Access layer of the Virtual Observatory. OpenSkyQuery (OSQ) allows cross matches between catalogs on registered nodes and supports the upload of lists of sources to be cross matched. This system utilizes the IVOA's nascent standard Astronomical Data Query Language (ADQL).
Abstract: The Sloan Digital Sky Survey (SDSS) science databasedescribes over 230 million objects and is over 1.6 TB insize. The SDSS Catalog Archive Server (CAS) providesseveral levels of query interface to the SDSS data via theSkyServer website. Most queries execute in seconds orminutes. However, some queries can take hours or days,either because they require non-index scans of the largesttables, or because they request very large result sets, orbecause they represent very complex aggregations...
We describe new and enhanced features for data access with the third data release (DR3) of the SDSS, particularly in the context of data intensive science and the VO. These include several enhancements to the CasJobs batch query workbench system, improved JPEG color images in the Visual Tools, extensive usage and traffic logging with harvesting of multiple remote site logs, enhancements to the HTM spatial library and a more versatile object crossid facility. We briefly describe these features and list future enhancements anticipated with SQLServer Yukon.
The Sloan Digital Sky Survey (SDSS) science database describes over 230 million objects and is over 1.6 TB in size. The SDSS Catalog Archive Server (CAS) provides several levels of query interface to the SDSS data via the SkyServer website. Most queries execute in seconds or minutes. However, some queries can take hours or days, either because they require non-index scans of the largest tables, or because they request very large result sets, or because they represent very complex aggregations of the data. These "monster queries" not only take a long time, they also affect response times for everyone else - one or more of them can clog the entire system. To ameliorate this problem, we developed a multiserver multiqueue batch job submission, execution, and tracking system for the CAS called CasJobs. The transfer of very large result sets from queries over the network is another serious problem. Statistics suggested that much of this data transfer is unnecessary; users would prefer to store results locally in order to allow further joins and filtering. To allow local analysis, a system was developed that gives users their own personal databases (MyDB) at the server side. Users may transfer data to their MyDB, and then perform further analysis before extracting it to their own machine. MyDB tables also provide a convenient way to share results of queries with collaborators without downloading them. CasJobs is built using SOAP XML Web services and has been in operation since May 2004.
The Sloan Digital Sky Survey (SDSS) has validated and made publicly available its Second Data Release. This data release consists of 3324 deg2 of five-band (ugriz) imaging data with photometry for over 88 million unique objects, 367,360 spectra of galaxies, quasars, stars, and calibrating blank sky patches selected over 2627 deg2 of this area, and tables of measured parameters from these data. The imaging data reach a depth of r ≈ 22.2 (95% completeness limit for point sources) and are photometrically and astrometrically calibrated to 2% rms and 100 mas rms per coordinate, respectively. The imaging data have all been processed through a new version of the SDSS imaging pipeline, in which the most important improvement since the last data release is fixing an error in the model fits to each object. The result is that model magnitudes are now a good proxy for point-spread function magnitudes for point sources, and Petrosian magnitudes for extended sources. The spectroscopy extends from 3800 to 9200 Å at a resolution of 2000. The spectroscopic software now repairs a systematic error in the radial velocities of certain types of stars and has substantially improved spectrophotometry. All data included in the SDSS Early Data Release and First Data Release are reprocessed with the improved pipelines and included in the Second Data Release. Further characteristics of the data are described, as are the data products themselves and the tools for accessing them.
The Sloan Digital Sky Survey science database is approaching 1TB in size. While the vast majority of queries normally execute in seconds or minutes, this prompt execution time can be disproportionately increased by a small fraction of queries that take hours or days to run either because they require non-index scans of the largest tables or because they request very large result sets. In response to this, a job submission and tracking system has been developed with multiple queues. The transfer of very large result sets from queries over the network is another serious problem. Statistics suggested that much of this data transfer is unnecessary; users would prefer to store results locally in order to allow further cross matching and filtering. To allow local analysis, a system was developed that gives users their own personal database (MYDB) at the portal site. Users may transfer data to their MYDB, then perform further analysis before extracting it to their own machine.
The Astronomical Data Query Language (ADQL) is a proposed standard query language for the interoperability of the International Virtual Observatory. The data servers in the International Virtual Observatory could be searched using an ADQL query. The servers would return VOTables as a result of the query.