This paper reports the results of a four- state collaboration-Illinois, New York, Texas, and Virginia-that uses Student Unit Record Database Systems that track students from high school into college. The goal is to determine whether it is possible to accurately predict whether individual students will not graduate using very early indicators available at college entry or during the first semester. Using similar statistical models across four state university systems, we identify individual students at greatest risk of non- completion quite accurately at early stages, allowing college staff to prioritize interventions and supports aimed at improving completion for those at greatest risk. Our logistic regression models rely on variables available to university administrators at student entry, including high school GPA, standardized test scores, parental income, remediation requirements, declared major, and college credits attempted in the first semester. Our models do not use gender, race, or ethnicity in determining probability of non- completion, making them useful for public university administrators. The fact that the same factors accurately predict graduation and non- completion in four very different state contexts suggests that similar dynamics are at play across the country. Our findings suggest that current commercial products that require extensive effort from faculty to input data on student progress, to act as an early warning system, may be unnecessary. More easily obtainable data can accurately predict students at risk of non- completion.
This paper reports the results of a four-state collaboration––Texas, New York, Virginia, and Illinois––that uses Student Unit Record Database Systems that track students from high school into college. The goal is to determine whether it is possible to accurately predict whether individual students will not graduate using very early indicators available at college entry or during the first semester. Using similar statistical models across four state university systems, we identify individual students at greatest risk of non-completion quite accurately at early stages, allowing college staff to prioritize interventions and supports aimed at improving completion for those at greatest risk. Our logistic regression models rely on variables available to university administrators at student entry, including high school GPA, standardized test scores, parental income, remediation requirements, declared major, and college credits attempted in the first semester. Our models do not use gender, race, or ethnicity in determining probability of non-completion, making them useful for public university administrators. The fact that the same factors accurately predict graduation and non-completion in four very different state contexts suggests that similar dynamics are at play across the country. Our findings suggest that current commercial products that require extensive effort from faculty to input data on student progress, to act as an early warning system, may be unnecessary. More easily obtainable data can accurately predict students at risk of non-completion.
Many undergraduates leave college without completing a degree or credential. Some researchers characterize this as a waste of the student's time because (they assert) college short of a degree does not yield any advantage in the labor market. Using data for an entire cohort of students graduating high school in Texas in one year, we compare the employment and earnings years later of those who do not go beyond high school with those who enter college but do not complete a credential. Using techniques that address selection bias, we find that students with "some college" are considerably more likely to be employed fifteen years after high school graduation and tend to earn significantly more than their counterparts who do not go to college. These benefits are found across student subgroups, with low-income students, women, and students of color generally experiencing the greatest improvements in labor outcomes from college attendance. While college dropouts do not fare as well as college graduates, incomplete college nevertheless functions for many as a stepping-stone into a better labor market position.
AbstractGlobal hydrological and land surface models are increasingly used for tracking terrestrial total water storage (TWS) dynamics, but the utility of existing models is hampered by conceptual and/or data uncertainties related to various underrepresented and unrepresented processes, such as groundwater storage. The gravity recovery and climate experiment (GRACE) satellite mission provided a valuable independent data source for tracking TWS at regional and continental scales. Strong interests exist in fusing GRACE data into global hydrological models to improve their predictive performance. Here we develop and apply deep convolutional neural network (CNN) models to learn the spatiotemporal patterns of mismatch between TWS anomalies (TWSA) derived from GRACE and those simulated by NOAH, a widely used land surface model. Once trained, our CNN models can be used to correct the NOAH‐simulated TWSA without requiring GRACE data, potentially filling the data gap between GRACE and its follow‐on mission, GRACE‐FO. Our methodology is demonstrated over India, which has experienced significant groundwater depletion in recent decades that is nevertheless not being captured by the NOAH model. Results show that the CNN models significantly improve the match with GRACE TWSA, achieving a country‐average correlation coefficient of 0.94 and Nash‐Sutcliff efficient of 0.87, or 14% and 52% improvement, respectively, over the original NOAH TWSA. At the local scale, the learned mismatch pattern correlates well with the observed in situ groundwater storage anomaly data for most parts of India, suggesting that deep learning models effectively compensate for the missing groundwater component in NOAH for this study region.
Texas Parks and Wildlife Department through U.S. Fish and Wildlife Service State Wildlife Grant Program, grant TX T-106-1 (CFDA# 15.634)
The Evaluating and Enhancing the eXtreme Digital Cyberinfrastructure for Maximum Usability and Science Impact project, known as the Technology Investigation Service (TIS), was a collaboration between University of Illinois at Urbana-Champaign National Center for Supercomputing Applications, Pittsburgh Supercomputing Center, The University of Texas at Austin Texas Advanced Computing Center, University of Tennessee National Institute for Computational Sciences, and University of Virginia which identified and evaluated potential technologies to close the gap between the XSEDE (http://www.xsede.org) service offerings and the needs of XSEDE users. This project was funded by the NSF Division of Advanced Cyberinfrastructure (award ACI 09-46505) in response to the "Technology Audit and Insertion Service" component of the "TeraGrid Phase III: eXtreme Digital Resources for Science and Engineering (XD)" solicitation (NSF 08-571). Over the project lifetime the two major goals of TIS were: 1) identifying, tracking, evaluating and making recommendations of new technologies to XSEDE for consideration of adoption and 2) raising awareness of TIS to XSEDE and other stakeholders to solicit their input on technologies for consideration for evaluation. In accomplishing the goals, the following four significant outcomes from TIS resulted: the development and deployment of the XSEDE Technology Evaluation Database; the development of the significantly improved XSEDE software search capability; the technology evaluation process; and the evaluations performed along with their corresponding technology adoption recommendations to XSEDE. This paper highlights the life-cycle of the TIS project, including lessons learned and project outcomes.
Poster presentation presented at the 2017 Texas Academy of Sciences annual meeting in Belton, Texas on March 4, 2017.
The growth in the capacity and capability of NAND Flash based storage systems have changed the face of data oriented computational systems. These systems have become both more capable and flexible in how they are used. With these changes comes both increased potential and user complexity. While many systems attempt to hide this complexity through the addition of more layers of storage caches, the design of the Wrangler system went a different route, choosing instead to build a simple yet flexible web based interface to allow users to easily configure this complex data computing system based on their service and software needs. This allows users to work in the environments best suited to their workflows while optimally utilizing the systems high performance and high capacity storage systems. This interface also allows users to schedule long term periods of reserved capacity, "data campaigns", for projects. Finally, the system has been designed to support the data storage and sharing capacities of the system to enable these key aspects of data research. We discuss the capabilities with respect to three already existing workflows on the system to highlight the diversity and flexibility provided by this environment to data researchers.
This paper has two objectives: 1) to describe the experimental and data collection methods for a large-scale smart grid deployment in Austin, Texas, and 2) to provide results based on those data. As of October 2012, the test bed was comprised of I) 250 homes concentrated in a single neighborhood all built after 2007, and 2) 160 homes distributed throughout Austin with ages ranging from 10 to 92 years old. This experiment includes 200 electric monitoring systems (15-s resolution), 211 electric monitoring systems (1-min), 182 gas meters (2-cubic foot), and 51 water meters (1 gallon) and many of the monitored homes also have energy audits and homeowner surveys. The test bed also includes 185 rooftop PV (photovoltaic) installations and 50 electric vehicles in the same neighborhood. Data streams were automated and gathered at a supercomputing facility at UT-Austin yielding 250 GB (2.95 x 10(9) records) of data in the first year. This paper describes the baseline study and monitoring methods, characterizes the study participants, and provides some first results about residential energy use. These results include a negative correlation between energy use and knowledge about energy as well as a possible positive correlation between energy use and some rebates. (C) 2013 Elsevier Ltd. All rights reserved.
Over the years, R has been adopted as a major data analysis and mining tool in many domain fields. As Big Data overwhelms those fields, the computational needs and workload of existing R solutions increases significantly. With recent hardware and software developments, it is possible to enable massive parallelism with existing R solutions with little to no modification. In this paper, we evaluated approaches to speed up R computations with the utilization of the Intel Math Kernel Library and automatic offloading to Intel Xeon Phi SE10P Co-processor. The testing workload includes a popular R benchmark and a practical application in health informatics. There are up to five times speedup gains from using MKL with a 16 cores without modification to the existing code for certain computing tasks. Offloading to Phi co-processor further improves the performance. The performance gains through parallelization increases as the data size increases, a promising result for adopting R for big data problem in the future.
This paper presents a data management scheme for the Pecan Street smart grid demonstration project in Austin, Texas. In this project, highly granular data with 15-second resolution on resource generation and consumption, including total consumption of electricity, water, and natural gas and solar generation, are collected for more than 100 homes. Furthermore, this testbed, see Figure 1, of homes represents the nation's highest density of rooftop solar PV and electric vehicles, and includes a substantial subset of homes that are highly instrumented with meters on up to 6 sub-circuits in addition to the whole-home meter. Consequently, this demonstration project generates a one-of-a-kind dataset with excellent temporal and geographic fidelity.One consequence of this extensive dataset is that there are hundreds of parallel data streams that need to be remotely (wirelessly) collected, filtered, processed, managed, stored and analyzed to be useful for researchers. Cumulatively, they represent 100s of gigabytes of data after just a few months of collection, which represents a formidable barrier to conducting research.In partnership with the Texas Advanced Computing Center (TACC), which is an NSF-sponsored cluster of supercomputers at UT-Austin, a data collection and management scheme has been developed. For storing the data, we have built a single column oriented database that so far has shown tremendous performance benefits. This paper shows the data schema, an example of MySQL query, and a developed program for rapid and automated[GRAPHICS]data extraction, analysis and display. We expect that the findings of this work will be beneficial to researchers interested in grid-scale data management.
The Texas Advanced Computing Center and the Institute for Classical Archaeology at the University of Texas at Austin developed a method that uses iRods rules and a Jython script to automate the extraction of metadata from digital archaeological data. The first step was to create a record-keeping system to classify the data. The record-keeping system employs file and directory hierarchy naming conventions designed specifically to maintain the relationship between the data objects and map the archaeological documentation process. The metadata implicit in the record-keeping system is automatically extracted upon ingest, combined with additional sources of metadata, and stored alongside the data in the iRods preservation environment. This method enables a more organized workflow for the researchers, helps them archive their data close to the moment of data creation, and avoids error prone manual metadata input. We describe the types of metadata extracted and provide technical details of the extraction process and storage of the data and metadata.
The requirements to support large-scale and complex research collections are growing at an accelerated pace. Considering the continuous evolution of the collections, their increasing sizes, the technologies supporting them, and the importance of adequate data management to long-term preservation, a team at the Texas Advanced Computing Center (TACC) developed a cyberinfrastructure to aid researchers in the creation, management, and curation of collections throughout the research lifecycle processes and beyond for access and long term preservation. Collections are maintained on a petabyte-scale data applications facility, and consulting services are available to address data curation needs. In this environment, researchers have the flexibility to build their collections without having to deal with details such as systems administration and hardware migration planning. The cyberinfrastructure facilitates the development of sustainable collections and a seamless transition through data gathering and curation, large-scale analysis, and collections dissemination and preservation.
GridPort is a toolkit for developing Web-based portals and applications for computational science on top of underlying distributed and grid computing infrastructure. GridPort aggregates grid services from popular grid software packages and provides additional grid capabilities, while presenting a simple, consistent API for portal and applications developers.
The paper describes the physical and mathematical fundamentals of dynamic gravimetry in liquid phase. First, a relation is developed to calculate the apparent weight of a vertical thin cylinder partially immersed Into a liquid. It emphasizes the magnitude of buoyancy forces and surface-tension forces (which act upon fine threads used to suspend a solid object completely immersed into a liquid). Next, the basic relation of dynamic gravimetry is derived to estimate the absolute mass of adsorbed material onto the surface of a submerged solid object. This relation encompasses, apart from the term related to the measured output signal of the balance, three correction terms which originate from the surface tension effects, buoyancy effect of the submerged object, and buoyancy effect due to the adsorbed material.The following sections address adsorption at the air-liquid interface as well as the molecular diffusion process (both being intermingled with concentration and density gradients in the system). Special attention is given to the errors connected with a series of relaxation processes in which the external menisci, the plate-beam system, and the surface tension are involved once their equilibrium conditions are disturbed. The time constants of these relaxation processes can not be exactly predicted. The paper concludes with an experimental illustration of dynamic gravimetry, with the adsorption of beta-lactoglobulin onto stainless steel plates. The validity of predictions based upon the physico-mathematical fundamentals is reconciled with the analytical and technical aspects and difficulties specific to dynamic gravimetry.