One of the two X-chromosomes in female mammals is epigenetically silenced in embryonic stem cells by X-chromosome inactivation. This creates a mosaic of cells expressing either the maternal or the paternal X allele. The X-chromosome inactivation ratio, the proportion of inactivated parental alleles, varies widely among individuals, representing the largest instance of epigenetic variability within mammalian populations. While various contributing factors to X-chromosome inactivation variability are recognized, namely stochastic and/or genetic effects, their relative contributions are poorly understood. This is due in part to limited cross-species analysis, making it difficult to distinguish between generalizable or species-specific mechanisms for X-chromosome inactivation ratio variability. To address this gap, we measure X-chromosome inactivation ratios in ten mammalian species (9531 individual samples), ranging from rodents to primates, and compare the strength of stochastic models or genetic factors for explaining X-chromosome inactivation variability. Our results demonstrate the embryonic stochasticity of X-chromosome inactivation is a general explanatory model for population X-chromosome inactivation variability in mammals, while genetic factors play a minor role. Stochastic or genetic effects can drive variability in X-chromosome inactivation (XCI) ratios among mammals, though their relative contributions are poorly understood. This study measures distributions of XCI ratios for 10 species and finds the embryonic stochasticity of XCI is a general explanatory model for mammalian XCI ratio variability.
Methods that provide specific, easy, and scalable experimental access to animal cell types and cell states will have broad applications in biology and medicine. CellREADR - Cell access through RNA sensing by Endogenous ADAR (adenosine deaminase acting on RNA), is a programmable RNA sensor-actuator technology that couples the detection of a cell-defining RNA to the translation of an effector protein to monitor and manipulate the cell. The CellREADR RNA device consists of a 5' sensor region complementary to a cellular RNA and a 3' payload coding region; payload translation is gated by the removal of a STOP codon in the sensor region upon base pairing with the cognate cellular RNA through an ADAR-mediated A- to-I editing mechanism ubiquitous to metazoan cells. CellREADR thus highlights the potential for RNA-based monitoring and manipulation of animal cells in ways that are simple, versatile, and generalizable across tissues and species. Here, we describe a detailed protocol for implementing CellREADR experiments in cell cultures and in animals. The procedure includes sensor and payload design, cloning, validation and characterization in mammalian cell cultures. The in vivo animal protocol focuses on AAV-based delivery of CellREADR using brain tissue as examples. We describe current best practices, various experimental controls and trouble-shooting steps. Beginning from sensor design, the validation of RNA sensor-actuators in cell cultures can be completed in 2-3 weeks. The construction of AAV vectors and their in vivo validation and characterization will take additional 6-8 weeks.
One of the two X chromosomes in female mammals is epigenetically silenced in embryonic stem cells by X chromosome inactivation (XCI). This creates a mosaic of cells expressing either the maternal or the paternal X allele. The XCI ratio, the proportion of inactivated parental alleles, varies widely among individuals, representing the largest instance of epigenetic variability within mammalian populations. While various contributing factors to XCI variability are recognized, namely stochastic and/or genetic effects, their relative contributions are poorly understood. This is due in part to limited cross-species analysis, making it difficult to distinguish between generalizable or species-specific mechanisms for XCI ratio variability. To address this gap, we measured XCI ratios in nine mammalian species (9,143 individual samples), ranging from rodents to primates, and compared the strength of stochastic models or genetic factors for explaining XCI variability. Our results demonstrate the embryonic stochasticity of XCI is a general explanatory model for population XCI variability in mammals, while genetic factors play a minor role.
X-chromosome inactivation (XCI) is a random, permanent, and developmentally early epigenetic event that occurs during mammalian embryogenesis. We harness these features to investigate characteristics of early lineage specification events during human development. We initially assess the consistency of X-inactivation and establish a robust set of XCI-escape genes. By analyzing variance in XCI ratios across tissues and indi-viduals, we find that XCI is shared across all tissues, suggesting that XCI is completed in the epiblast (in at least 6-16 cells) prior to specification of the germ layers. Additionally, we exploit tissue-specific variability to characterize the number of cells present during tissue-lineage commitment, ranging from approximately 20 cells in liver and whole blood tissues to 80 cells in brain tissues. By investigating the variability of XCI ratios using adult tissue, we characterize embryonic features of human XCI and lineage specification that are other-wise difficult to ascertain experimentally.
is low cost scrubbing of all gas contaminants from indoor air using novel, efficient and regenerable sorbent materials. The total solution, labeled HLR (HVAC Load Reduction) is implemented by means of compact modules that can be retrofitted onto virtually any building, and utilizes novel sorbents that have been developed and brought to scale production by enVerid. Although the technology had been deployed and demonstrated, the US commercial building and HVAC markets, famously conservative with respect to innovation and new technologies, had yet to embrace the idea or the product. The overarching goal of the project was to help overcome market hesitancy and enable widespread adoption, by creating a critical mass of success stories across different regions and building types, all under oversight and verification of the DOE and NREL, to ensure credibility and accuracy. The project, overall, was a resounding success. Several installations were deployed and tested comprehensively for the impact on both energy consumption and indoor air quality, and results were not only excellent, but quantitatively in line with expectations. A total number of eight (8) installations were attempted, with 2 aborted due to building issues. Of the six successful installations, however, complete M&V was performed by NREL on 2 (Miami and NYC) and partial analysis on a third site (ArcBest), and these three, along with their results, were therefore presented from a financial and reporting standpoint. Although the number of case studies fell short of the original target of 10, we nevertheless had excellent demonstrations in vastly different climate zones: Southeast, Northeast and Midwest, and in very different building types. This included a modern high-rise office building in New York City and a University Campus health center in Miami, and a mid-rise high-density office in Arkansas. The cooling load reduction was shown to be significant, averaging 37% during the final cooling season in the Miami site. In the NYC site, the cooling energy savings were generally less, at around 12%, as expected. While this is a lower number for average/total savings, two comments are important to note: (i) the outside airflow (OA) reduction was not consistent during the data collection period, at least in part understating the attainable savings through incorrect measurement of the “high energy” reference state; and (ii) the peak impact in NYC is almost certainly as high as Miami, with critical ramifications for peak load reduction and equipment downsizing ability. Importantly, the measured energy savings in all locations were related to the outdoor climate conditions and the actual achievable reduction in outside air, very much in line with models and calculations; this leaves us with increased confidence in our understanding of the interaction between the technology and the building, and our ability to project the impact of this innovation in the future, through broad adoption. Simply put: the energy savings and load reduction impact of the HLR solution has been confirmed quantitatively. The results also confirmed the air quality objective. This may be even more significant, in that air quality is more complex and harder to measure and quantify, and the performance of the HLR as an air scrubber is the linchpin of the entire solution. In all sites, indoor air quality met or exceeded the required parameters, in terms of achieving the target (low) levels of indoor pollutants. Furthermore, the concentrations of outdoor sourced pollutants were markedly lower, as expected, mostly due to the reduction in outside air intake although in part also due to the scrubbing. The success was not immediate nor easy, and early challenges with the systems, especially (but not only) around controls and communication, caused delays in commissioning and undermined some of our M&V schedules. The good news is no fundamental issues were encountered and the number of such glitches was declining quickly by the end of the project. Other challenges had to do with coordinating the control of the building OA dampers, which the HLR needs to have in order to deliver the load reduction, but which the facility manager controls regularly and independently. This is not necessarily a problem for normal HLR operation but interferes with accurate M&V, where switching OA settings back and forth on a precise timetable is essential for clean data. The project provided valuable experience in retrofitting buildings that had been originally designed to operate with less recirculation and more outside air, and some key learnings related to the cyclical regeneration of the sorbents in the operational environment of the building. Some of these learnings gave already been incorporated into ongoing improvements in the design of the HLR module as well as the software that is used to control its operation. The cost and time of retrofit installations is much better understood; so is the test and balance procedure of the building that determines the available range of outside air reduction, which in turn determines the amount of load reduced. We had also implemented at least one annual cycle of sorbent cartridge replacement in the field, after which we tested the used sorbent to assess the rate of degradation of sorbent potency – an important factor affecting the long-term economics of HLR deployments. A notable shortfall of the project, other than the fact that the number of sites was below target, was that we had not (yet) completed testing to measure and verify (M&V) the winter energy savings. This is primarily due to late start on the NYC sites (Miami does not portend any heating savings associated with reduced outside air), combined with the issues of OA damper control. It is by no means suggested that such savings cannot be captured, only that the M&V with respect to winter in NYC has been incomplete. The value of this project and its interest to the general public are hard to overstate. The case studies created from Miami and especially from NYC have, without a doubt, opened up the interest of the broad HVAC industry to HLR in the past year. Major manufacturer reps have signed up to represent and promote the product, and dozens of top tier engineering consultants have introduced HLR as a base of design in upcoming construction and renovation projects – precisely the overarching goal of this project. These would not have happened without the highly visible installation in NYC and its ultimately satisfied customer. Furthermore, the DOE itself issued a report in late 2017 analyzing hundreds of candidate technologies for reducing energy consumption in HVAC; remarkably, the budding HLR solution ranked among the top 3(!); no doubt, its status as a viable and significant contender on this list is informed by the success of these demonstrations. And furthermore, in the wake of these demos, several major energy utilities, led by ConEdison in NYC, have endorsed the HLR solution and now routinely provide substantial cash rebates for customers choosing to incorporate the HLR in their buildings and reap the benefit of the reduced peak load as well as reduced cumulative energy demand. Last but not least, in January of 2019, as this report was being written, the HLR module won the AHR Expo Innovation Award for Green Buildings and, remarkably, the prestigious and only Product of the Year 2019, being selected by an independent professional panel of judges from AHR and ASHRAE. Given the confirmation of the enormous energy saving potential of HLR in widescale deployment, as well as ancillary but significant benefits in terms of equipment downsizing and indoor air quality, the benefit of this effort to the public speaks for itself. Accelerating adoption of HLR, without dependence on subsidies and driven by the private sector, is well under way. There is still much work to be done in terms of educating the HVAC and construction ecosystems and improving the product’s robustness and simplicity, to further facilitate the adoption of this technology and the capture of its potential value. But the American public will greatly benefit both directly and indirectly in the immediate future and beyond.
X-chromosome inactivation (XCI) is a random, permanent, and developmentally early epigenetic event that occurs during mammalian embryogenesis. We harness these features of XCI to investigate characteristics of early lineage specification events during human development. We initially assess the consistency of X-inactivation and establish a robust set of XCI-escape genes. By analyzing variance in XCI ratios across tissues and individuals, we find that XCI is completed prior to tissue specification and at a time when 6-16 cells are fated for all tissue lineages. Additionally, we exploit tissue specific variability to characterize the number of cells present at the time of each tissue’s lineage commitment, ranging from approximately 20 cells in liver and whole blood tissues to 80 cells in brain tissues. By investigating variance of XCI ratios using adult tissue, we resolve key features of human development otherwise difficult to ascertain experimentally and develop scalable methods easily applicable to future data.
BNL SDCC (Scientific Data and Computing Center) recently deployed a centralized identity management solution to support Single Sign On (SSO) authentication across multiple IT systems. The system supports federated login access via CILogon and InCommon and multi-factor authentication (MFA) to meet security standards for various application and services such as Jupyterhub / Invenio that are provided to the SDCC user community. CoManage (cloud-based) and FreeIPA / Keycloak (local) are utilized to provided complex authorization for authenticated users. This talk will focus on technical overviews and strategies to tackle the challenges/obstacles in our facility.
The Production and Distributed Analysis (PanDA) system has been successfully used in the ATLAS experiment as a data-driven workload management system. The PanDA system has proven to be capable of operating at the Large Hadron Collider data processing scale over the last decade including the Run 1 and Run 2 data taking periods. PanDA was originally designed to be weakly coupled with the WLCG processing resources. Lately the system is revealing the difficulties to optimally integrate and exploit new resource types such as HPC and preemptible cloud resources with instant spin-up, and new workflows such as the event service, because their intrinsic nature and requirements are quite different from that of traditional grid resources. Therefore, a new component, Harvester, has been developed to mediate the control and information flow between PanDA and the resources, in order to enable more intelligent workload management and dynamic resource provisioning based on detailed knowledge of resource capabilities and their real-time state. Harvester has been designed around a modular structure to separate core functions and resource specific plugins, simplifying the operation with heterogeneous resources and providing a uniform monitoring view. This paper will give an overview of the Harvester architecture, current status with various resources, and future plans.
HEPCloud is rapidly becoming the primary system for provisioning compute resources for all Fermilab-affiliated experiments. In order to reliably meet the peak demands of the next generation of High Energy Physics experiments, Fermilab must plan to elastically expand its computational capabilities to cover the forecasted need. Commercial cloud and allocation-based High Performance Computing (HPC) resources both have explicit and implicit costs that must be considered when deciding when to provision these resources, and at which scale. In order to support such provisioning in a manner consistent with organizational business rules and budget constraints, we have developed a modular intelligent decision support system (IDSS) to aid in the automatic provisioning of resources spanning multiple cloud providers, multiple HPC centers, and grid computing federations. In this paper, we discuss the goals and architecture of the HEPCloud Facility, the architecture of the IDSS, and our early experience in using the IDSS for automated facility expansion both at Fermi and Brookhaven National Laboratory.
A traditional HPC computing facility provides a large amount of computing power but has a fixed environment designed to satisfy local needs. This makes it very challenging for users to deploy complex applications that span multiple sites and require specific application software, scheduling middleware, or sharing policies. The DOE-funded VC3 project aims to address these challenges by making it possible for researchers to easily aggregate and share resources, install custom software environments, and deploy clustering frameworks across multiple HPC facilities through the concept of "virtual clusters". This paper presents the design, implementation, and initial experience with our prototype self-service VC3 platform which automates deployment of cluster frameworks across diverse computing facilities. To create a virtual cluster, the VC3 platform materializes a custom head node in a secure private cloud, specifies a choice of scheduling middleware, then allocates resources from the remote facilities where the desired software and clustering framework is installed in user space. As resources become available from scheduled nodes from individual clusters, the research team simply sees a private cluster they can access directly or share with collaborators, such as a science gateway community. We discuss how this service can be used by research collaborations requiring shared resources, specific middleware frameworks, and complex applications and workflows in the areas of astrophysics, bioinformatics and high energy physics.
Throughout the first half of LHC Run 2, ATLAS cloud computing has undergone a period of consolidation, characterized by building upon previously established systems, with the aim of reducing operational effort, improving robustness, and reaching higher scale. This paper describes the current state of ATLAS cloud computing. Cloud activities are converging on a common contextualization approach for virtual machines, and cloud resources are sharing monitoring and service discovery components. We describe the integration of Vacuum resources, streamlined usage of the Simulation at Point 1 cloud for offline processing, extreme scaling on Amazon compute resources, and procurement of commercial cloud capacity in Europe. Finally, building on the previously established monitoring infrastructure, we have deployed a real-time monitoring and alerting platform which coalesces data from multiple sources, provides flexible visualization via customizable dashboards, and issues alerts and carries out corrective actions in response to problems.
Over the past few years, Grid Computing technologies have reached a high level of maturity. One key aspect of this success has been the development and adoption of newer Compute Elements to interface the external Grid users with local batch systems. These new Compute Elements allow for better handling of jobs requirements and a more precise management of diverse local resources.
Continued growth in public cloud and HPC resources is on track to exceed the dedicated resources available for ATLAS on the WLCG. Examples of such platforms are Amazon AWS EC2 Spot Instances, Edison Cray XC30 supercomputer, backfill at Tier 2 and Tier 3 sites, opportunistic resources at the Open Science Grid (OSG), and ATLAS High Level Trigger farm between the data taking periods. Because of specific aspects of opportunistic resources such as preemptive job scheduling and data I/O, their efficient usage requires workflow innovations provided by the ATLAS Event Service. Thanks to the finer granularity of the Event Service data processing workflow, the opportunistic resources are used more efficiently. We report on our progress in scaling opportunistic resource usage to double-digit levels in ATLAS production.
After a scheduled maintenance and upgrade period, the world’s largest and most powerful machine – the Large Hadron Collider(LHC) – is about to enter its second run at unprecedented energies. In order to exploit the scientific potential of the machine, the experiments at the LHC face computational challenges with enormous data volumes that need to be analysed by thousand of physics users and compared to simulated data. Given diverse funding constraints, the computational resources for the LHC have been deployed in a worldwide mesh of data centres, connected to each other through Grid technologies.The PanDA (Production and Distributed Analysis) system was developed in 2005 for the ATLAS experiment on top of this heterogeneous infrastructure to seamlessly integrate the computational resources and give the users the feeling of a unique system. Since its origins, PanDA has evolved together with upcoming computing paradigms in and outside HEP, such as changes in the networking model, Cloud Computing and HPC. It is currently running steadily up to 200 thousand simultaneous cores (limited by the available resources for ATLAS), up to two million aggregated jobs per day and processes over an exabyte of data per year. The success of PanDA in ATLAS is triggering the widespread adoption and testing by other experiments. In this contribution we will give an overview of the PanDA components and focus on the new features and upcoming challenges that are relevant to the next decade of distributed computing workload management using PanDA.
The Large Hadron Collider(LHC) is the world's largest and most powerful machine. It started operating in 2009 with a scientific program foreseen to extend over the next coming decades at increasing energies and luminosities to maximise the discovery potential. During Run1 (2009- 2013), the Worldwide LHC Computing Grid (WLCG) successfully delivered all the necessary computing resources, which made the discovery of the Higgs Boson possible. Looking ahead, it is forecasted that increased luminosities will extrapolate to a multiplicity in the storage and processing costs, which is not reflected in a corresponding funding growth of the WLCG. ATLAS, one of the four experiments at the LHC, is therefore leading an upgrade program to evolve their software and computing model to make the best possible usage of available resources, and also leverage on upcoming state of the art computing paradigms that could make important resource contributions. These proceedings will give an insight into the accompanying work in PanDA, ATLAS’ workload management system. PanDA has implemented event level bookkeeping and dynamic generation of jobs with tailored lengths, in order to integrate and optimise the usage of oppor- tunistic resources, e.g. Cloud Computing or High Performance Computing (HPC). In conjunc- tion, the Event Service has been developed as a way to manage fine grained jobs and its outputs. Usage examples on some of the leading commercial and research infrastructures will be given. In addition, we will describe the work on further exploiting the current network capabilities by allowing remote data access and reducing regional boundaries.
The ATLAS experiment at the LHC has successfully incorporated cloud computing technology and cloud resources into its primarily grid-based model of distributed computing. Cloud R&D activities continue to mature and transition into stable production systems, while ongoing evolutionary changes are still needed to adapt and refine the approaches used, in response to changes in prevailing cloud technology. In addition, completely new developments are needed to handle emerging requirements.This paper describes the overall evolution of cloud computing in ATLAS. The current status of the virtual machine (VM) management systems used for harnessing Infrastructure as a Service resources are discussed. Monitoring and accounting systems tailored for clouds are needed to complete the integration of cloud resources within ATLAS' distributed computing framework. We are developing and deploying new solutions to address the challenge of operation in a geographically distributed multi-cloud scenario, including a system for managing VM images across multiple clouds, a system for dynamic location-based discovery of caching proxy servers, and the usage of a data federation to unify the worldwide grid of storage elements into a single namespace and access point. The usage of the experiment's high level trigger farm for Monte Carlo production, in a specialized cloud environment, is presented. Finally, we evaluate and compare the performance of commercial clouds using several benchmarks.
The computing model of the ATLAS experiment was designed around the concept of grid computing and, since the start of data taking, this model has proven very successful. However, new cloud computing technologies bring attractive features to improve the operations and elasticity of scientific distributed computing. ATLAS sees grid and cloud computing as complementary technologies that will coexist at different levels of resource abstraction, and two years ago created an R&D working group to investigate the different integration scenarios. The ATLAS Cloud Computing R&D has been able to demonstrate the feasibility of offloading work from grid to cloud sites and, as of today, is able to integrate transparently various cloud resources into the PanDA workload management system. The ATLAS Cloud Computing R&D is operating various PanDA queues on private and public resources and has provided several hundred thousand CPU days to the experiment. As a result, the ATLAS Cloud Computing R&D group has gained a significant insight into the cloud computing landscape and has identified points that still need to be addressed in order to fully utilize this technology. This contribution will explain the cloud integration models that are being evaluated and will discuss ATLAS' learning during the collaboration with leading commercial and academic cloud providers.
The Production and Distributed Analysis system (PanDA) has been in use in the ATLAS Experiment since 2005. It uses a sophisticated pilot system to execute submitted jobs on the worker nodes. While originally designed for ATLAS, the PanDA Pilot has recently been refactored to facilitate use outside of ATLAS. Experiments are now handled as plug-ins such that a new PanDA Pilot user only has to implement a set of prototyped methods in the plug-in classes, and provide a script that configures and runs the experiment-specific payload. We will give an overview of the Next Generation PanDA Pilot system and will present major features and recent improvements including live user payload debugging, data access via the Federated XRootD system, stage-out to alternative storage elements, support for the new ATLAS DDM system (Rucio), and an improved integration with glExec, as well as a description of the experiment-specific plug-in classes. The performance of the pilot system in processing LHC data on the OSG, LCG and Nordugrid infrastructures used by ATLAS will also be presented. We will describe plans for future development on the time scale of the next few years.
The RHIC and ATLAS Computing Facility (RACF) at Brookhaven Lab is a dedicated data center serving the needs of the RHIC and US ATLAS community. Since it began operations in the mid-1990's, it has operated continuously with few unplanned downtimes. In the past 15 months, Brookhaven Lab has been affected by two hurricanes and a record-breaking snowstorm. In this presentation, we discuss lessons learned regarding (natural or man-made) disaster preparedness, operational continuity, remote access and safety protocols, including overall operational procedures developed as a result of these recent events.
The Open Science Grid encourages the concept of software portability: a user's scientific application should be able to run at as many sites as possible. It is necessary to provide a mechanism for OSG Virtual Organizations to install software at sites. Since its initial release, the OSG Compute Element has provided an application software installation directory to Virtual Organizations, where they can create their own sub-directory, install software into that sub-directory, and have the directory shared on the worker nodes at that site. The current model has shortcomings with regard to permissions, policies, versioning, and the lack of a unified, collective procedure or toolset for deploying software across all sites. Therefore, a new mechanism for data and software distributing is desirable. The architecture for the OSG Application Software Installation Service (OASIS) is a server-client model: the software and data are installed only once in a single place, and are automatically distributed to all client sites simultaneously. Central file distribution offers other advantages, including server-side authentication and authorization, activity records, quota management, data validation and inspection, and well-defined versioning and deletion policies. The architecture, as well as a complete analysis of the current implementation, will be described in this paper.
Yuri Demchenko合作论文数Fraunhofer Institut SCAI, 53754 Sankt Augustin, GermanyPoznan Supercomputing and Networking Center, Noskowskiego 12/14 , 61-704 Poznan, Poland2