The ATLAS experiment at the Large Hadron Collider (LHC), located in Geneva, will collect 3 to 4 petabytes or PB (1015 bytes) of data for each year of its operation, when fully commissioned. Secondary data sets resulting from event reconstruction, reprocessing and calibration will result in an additional 2.5 PB for each year of data taking. Simulated data sets require also significant resources as well nearing 1 PB per year. The data will be distributed worldwide to ten Tier-1 computing centres within the Worldwide LHC computing grid (WLCG) that will operate around the clock. One of these centres is hosted at TRIUMF, Canada's National Laboratory for Particle and Nuclear Physics, located in Vancouver, BC. By the year 2010, the storage capacity at TRIUMF will consist of about 3 Petabyte of disk storage, and 2 PB of tape storage. At present, the disk capacity installed is 750 terabytes or TB (1012 bytes) while the tape capacity is 560 TB, both using state of the art technology. dCache from www.dcache.org is used to manage the entire storage in order to provide a common file namespace. It is a highly scalable solution and highly configurable. In this paper we will describe and review the storage infrastructure and configuration currently in place at the Tier-1 centre at TRIUMF for both disk and tape as well as the management software and tools that have been developed.
Networking at an ATLAS Tier-1 (Tl) facility is a demanding aspect which is vital to the overall performance and efficiency of the facility. External connectivity of the facility to other tiers of the large hadron collider optical private network (LHCOPN) is largely via dedicated lightpaths as required to meet memorandum of understanding (MOU) commitments. Our primary dedicated link to CERN has an independent, although smaller capacity, dedicated backup link for redundancy. Dedicated lightpaths to the Canadian Tier-2 facilities, and the international partner Tier-1 facilities failover to national and international research networks in the event of failure. The distance between TRIUMF and CERN, and even TRIUMF to some of it's Tier-2 facilities in Canada is thousands of kilometers. Transferring data at the hundreds of terabytes of data level (per year) over such distances and complex networks requires both dedicated bandwidth and network resiliency. Failure scenario, including both fail- over and fail-back, must be handled efficiently and for the most part automatically. Although modern network routing protocols handle this well, monitoring processes become key to management of the infrastructure as the complexity of our Tier-1 site connectivity grows. Internal networking efficiency is also vital to the ATLAS computing model. Large data sets are moved from onsite storage to local scratch disks before analysis, and proper network scaling is vital for efficient use of compute nodes. Although 10 Gigabit network infrastructure (routers) is well established, server 10 Gigabit network components and drivers are not as mature as their I Gigabit counterparts, and resource issues have been observed during extreme load testing. In this paper, internal facility networking, as well as external connectivity issues will be discussed.
In preparation for the startup of the LHC at CERN, the ATLAS collaboration is performing computing challenges of increasing complexity and scale. Monte Carlo production is ongoing with some 10,000 concurrent jobs worldwide, including at Canadian collaborating institutes. TRIUMF hosts one of the 10 ATLAS 'Tier-1' data and computing centres, where 270 TB of data per year will be transferred, by high performance dedicated networks, processed and redistributed for analysis. We discuss recent large scale tests of all aspects of middleware, hardware infrastructure and procedures to support the 24 times 7 operation of the Tier-1 centre, required for LHC running.