Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · Correlated failures · OSDI 2010

Google measured a disk MTTF of 10 to 50 years but a node MTTF of 4.3 months; 37% of node failures came in bursts

Ford et al., OSDI 2010, studied tens of Google storage cells over one year. Node MTTF was 4.3 months against 10 to 50 years for a disk, and 37% of failures were part of a burst of at least 2 nodes. Cross-check rows come from NetApp (FAST 2020) and HPC1 (FAST 2007) and show clustered drive replacements.

Download CSV14 rows · 7 columns · CSV

File /data/datasets/storage-availability-correlated-failures.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Ford et al. (OSDI 2010) and cross-check rows from Maneas et al. (FAST 2020) and Schroeder and Gibson (FAST 2007). It does not read chart values.

Limits

Table

14 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-10.
MeasureValue lowValue highUnitScopeNoteCitation
Storage nodes per cell studied10007000nodes per cellTens of Google storage cells, one yearPrinted as 1000 to 7000 nodes in each cell.ford-osdi10-pdf
Node unavailability events lasting longer than 15 minutesnot stated10percent of eventsSame cellsPrinted as less than 10%; the paper then studies only events of 15 minutes or longer.ford-osdi10-pdf
Disk mean time to failure1050yearsGoogle cells; counts software and hardware causesTable 2 range; disk failures are permanent.ford-osdi10-pdf
Node mean time to failurenot stated4.3monthsSame cellsTable 2; most node failures are transient.ford-osdi10-pdf
Rack mean time to failurenot stated10.2yearsSame cellsTable 2.ford-osdi10-pdf
Failures that are part of a burst of at least 2 nodesnot stated37percent of failuresBursts defined with a 120 second windowPrinted as 37%; the authors estimate that close to 37% are truly correlated.ford-osdi10-pdf
Random failures wrongly merged into a burstnot stated8.0percent of non-correlated failuresSame window; 0.068% for a burst of at least 10 nodesUpper estimate of mis-clustering.ford-osdi10-pdf
Older data blocks that fail a checksum on scrubbing0.00000010.000001fraction of older blocksGFS background scrubbingPrinted as 1 in 10^6 to 10^7; concentrated on a small number of disks.ford-osdi10-pdf
Gain in stripe availability from a 10% cut in the disk failure ratenot stated1.5percent (upper bound)Model with R = 3 replicationPrinted as less than 1.5%.ford-osdi10-pdf
Gain in data availability from a 10% cut in the node failure ratenot stated18percentSame modelMore than 12 times the disk figure.ford-osdi10-pdf
Chance of a drive replacement in a random weeknot stated0.0504percentNetApp RAID groups, about 1.4 million SSDsBaseline for the next row.maneas-fast20-pdf
Chance of a replacement within a week of a previous one in the same RAID groupnot stated9.39percentSame fleetMore than 180 times the baseline; 52% of consecutive replacements fall within a week.maneas-fast20-pdf
Correlation of disk replacements in consecutive weeksnot stated0.72correlation coefficientHPC1 system, whole lifetime0.79 for consecutive months; 0.4 to 0.8 when one year is used.schroeder-fast07-pdf
Spread in expected weekly replacements after a quiet versus a busy weeknot stated9timesHPC1 systemPrinted as a factor of 9 between the first and third bucket.schroeder-fast07-pdf

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'less than' it is the stated bound.
unit (string)
Unit of the value columns: nodes per cell, percent, years, months, fraction, times, or correlation coefficient.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Ford et al. (OSDI 2010), Maneas et al. (FAST 2020), and Schroeder and Gibson (FAST 2007), with credit. Not a Creative Commons license. The USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.

Sources

  1. Availability in Globally Distributed Storage Systems, Ford, Labelle, Popovici, Stokely, Truong, Barroso, Grimes, and Quinlan, USENIX OSDI 2010 (USENIX PDF) (accessed 2026-10-10)
  2. A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements) (accessed 2026-10-10)
  3. Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 5.2 on correlations) (accessed 2026-10-10)