The present invention provides a method and a system for distributing data in a distributed storage system. The method comprises the following steps of: receiving a file on data processing hardware; dividing the file received by the data processing hardware into blocks, wherein the blocks are data blocks and non-data blocks; grouping the blocks into a group by the data processing hardware; determining, by data processing hardware, a distribution of blocks of a group among storage devices of a distributed storage system based on a maintenance hierarchy of the distributed storage system, the maintenance hierarchy comprising hierarchical maintenance levels and maintenance domains, each maintenance domain having an active state or an inactive state, each storage device being associated with atleast one maintenance domain; and distributing group blocks to the storage devices by the data processing hardware according to the determined distribution, the group blocks being distributed among the plurality of maintenance domains to maintain an ability to reconstruct the blocks of the group while the maintenance domains are inactive.
PROBLEM TO BE SOLVED: To provide a data distribution method that allows a user to use a distributed storage system even if a part of the distributed storage system is in a maintenance period.SOLUTION: The method for distributing data includes a step of receiving a file in a non-transitory memory (802), a step of splitting the received file into chunks using a computer processor which communicates with the non-transitory memory (804), and a step of distributing the chunks to storage devices of a distributed storage system based on a maintenance hierarchy of the distributed storage system (806). The maintenance hierarchy includes maintenance units, each of which has an active state and an inactive state. Further, the individual storage devices are associated with the maintenance units. In order to maintain file accessibility of files when the maintenance units are in an inactive state, the chunks are distributed to a plurality of maintenance units.SELECTED DRAWING: Figure 8
A method of prioritizing data for recovery in a distributed storage system includes, for each stripe of a file having chunks, determining whether the stripe comprises high-availability chunks or low-availability chunks and determining an effective redundancy value for each stripe. The effective redundancy value is based on the chunks and any system domains associated with the corresponding stripe. The distributed storage system has a system hierarchy including system domains. Chunks of a stripe associated with a system domain in an active state are accessible, whereas chunks of a stripe associated with a system domain in an inactive state are inaccessible. The method also includes reconstructing substantially immediately inaccessible, high-availability chunks having an effective redundancy value less than a threshold effective redundancy value and reconstructing the inaccessible low-availability and other inaccessible high-availability chunks, after a threshold period of time.
Ernst W. Mayr合作论文数Lehrstuhl fur Effiziente Algorithmen
Institut fur Informatik
Technische Universitat Munchen1