Dataset · Correlated failures · OSDI 2010
Google measured a disk MTTF of 10 to 50 years but a node MTTF of 4.3 months; 37% of node failures came in bursts
Ford et al., OSDI 2010, studied tens of Google storage cells over one year. Node MTTF was 4.3 months against 10 to 50 years for a disk, and 37% of failures were part of a burst of at least 2 nodes. Cross-check rows come from NetApp (FAST 2020) and HPC1 (FAST 2007) and show clustered drive replacements.
Download CSV14 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Ford et al. (OSDI 2010) and cross-check rows from Maneas et al. (FAST 2020) and Schroeder and Gibson (FAST 2007). It does not read chart values.
Limits
- One operator's cells and software; figures describe nodes and racks as well as disks.
- Disk MTTF is a printed range of 10 to 50 years, not a vendor rating.
- The 10% sensitivity results come from the authors' model with 3-way replication.
- The 37% burst share depends on a 120 second window.
- The cross-check rows are other fleets with other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| Storage nodes per cell studied | 1000 | 7000 | nodes per cell | Tens of Google storage cells, one year | Printed as 1000 to 7000 nodes in each cell. | ford-osdi10-pdf |
| Node unavailability events lasting longer than 15 minutes | not stated | 10 | percent of events | Same cells | Printed as less than 10%; the paper then studies only events of 15 minutes or longer. | ford-osdi10-pdf |
| Disk mean time to failure | 10 | 50 | years | Google cells; counts software and hardware causes | Table 2 range; disk failures are permanent. | ford-osdi10-pdf |
| Node mean time to failure | not stated | 4.3 | months | Same cells | Table 2; most node failures are transient. | ford-osdi10-pdf |
| Rack mean time to failure | not stated | 10.2 | years | Same cells | Table 2. | ford-osdi10-pdf |
| Failures that are part of a burst of at least 2 nodes | not stated | 37 | percent of failures | Bursts defined with a 120 second window | Printed as 37%; the authors estimate that close to 37% are truly correlated. | ford-osdi10-pdf |
| Random failures wrongly merged into a burst | not stated | 8.0 | percent of non-correlated failures | Same window; 0.068% for a burst of at least 10 nodes | Upper estimate of mis-clustering. | ford-osdi10-pdf |
| Older data blocks that fail a checksum on scrubbing | 0.0000001 | 0.000001 | fraction of older blocks | GFS background scrubbing | Printed as 1 in 10^6 to 10^7; concentrated on a small number of disks. | ford-osdi10-pdf |
| Gain in stripe availability from a 10% cut in the disk failure rate | not stated | 1.5 | percent (upper bound) | Model with R = 3 replication | Printed as less than 1.5%. | ford-osdi10-pdf |
| Gain in data availability from a 10% cut in the node failure rate | not stated | 18 | percent | Same model | More than 12 times the disk figure. | ford-osdi10-pdf |
| Chance of a drive replacement in a random week | not stated | 0.0504 | percent | NetApp RAID groups, about 1.4 million SSDs | Baseline for the next row. | maneas-fast20-pdf |
| Chance of a replacement within a week of a previous one in the same RAID group | not stated | 9.39 | percent | Same fleet | More than 180 times the baseline; 52% of consecutive replacements fall within a week. | maneas-fast20-pdf |
| Correlation of disk replacements in consecutive weeks | not stated | 0.72 | correlation coefficient | HPC1 system, whole lifetime | 0.79 for consecutive months; 0.4 to 0.8 when one year is used. | schroeder-fast07-pdf |
| Spread in expected weekly replacements after a quiet versus a busy week | not stated | 9 | times | HPC1 system | Printed as a factor of 9 between the first and third bucket. | schroeder-fast07-pdf |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'less than' it is the stated bound.
- unit
- Unit of the value columns: nodes per cell, percent, years, months, fraction, times, or correlation coefficient.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Ford et al. (OSDI 2010), Maneas et al. (FAST 2020), and Schroeder and Gibson (FAST 2007), with credit. Not a Creative Commons license. The USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- Availability in Globally Distributed Storage Systems, Ford, Labelle, Popovici, Stokely, Truong, Barroso, Grimes, and Quinlan, USENIX OSDI 2010 (USENIX PDF)
- A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements)
- Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 5.2 on correlations)