Dataset · SSD failures · SYSTOR 2016
In an 18-month Berkeley study of 368 disks, 7 data disks failed against 13 enclosure backplane failures; disks were 90% of parts but about 4% of errors
Talagala and Patterson (UC Berkeley, 1999) logged 18 months of a prototype with 368 SCSI data disks. 7 data disks failed (1.9%), against 13 of 46 enclosures, and disks were about 90% of parts but around 4% of errors. Cross-check rows come from Schroeder and Gibson (FAST 2007) and Ford et al. (OSDI 2010).
Download CSV15 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.
Limits
- One small prototype built from 1996 parts; says little about modern drives.
- Counts are small: 7 disk failures and 1 failure for several other types.
- Failures are replacements over 18 months, not annual rates.
- The error table and the text disagree slightly on data disk errors and machine counts.
- The cross-check rows are other systems with other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| SCSI data disks in the prototype | not stated | 368 | disks | 20 PCs, 3.2 TB, 8.4 GB disks, hardware from 1996 | Disks were about 90% of the components. | talagala-ucb99-pdf |
| Length of the failure record | not stated | 18 | months | Prototype operation | Error logs cover the last 6 months. | talagala-ucb99-pdf |
| Components replaced in the failure record | not stated | 34 | components | All component types | Printed as 34 absolute failures, nearly two a month. | talagala-ucb99-pdf |
| SCSI data disks replaced | not stated | 1.9 | percent of 368 disks in 18 months | 7 of 368 disks | Lowest percentage of any component that failed. | talagala-ucb99-pdf |
| Disk enclosure backplane failures | not stated | 28.3 | percent of 46 enclosures in 18 months | 13 of 46 enclosures | Backplane integrity failures on the SCSI bus. | talagala-ucb99-pdf |
| IDE system disks replaced | not stated | 25.0 | percent of 24 disks in 18 months | 6 of 24 disks | Mounted in ordinary PC chassis, not the cooled enclosures. | talagala-ucb99-pdf |
| Enclosure power supplies replaced | not stated | 3.26 | percent of 92 supplies in 18 months | 3 of 92 supplies | Each enclosure had two supplies, so a failure did not take it offline. | talagala-ucb99-pdf |
| Error instances in the logs | not stated | 688 | error instances | 6 months of logs, 16 or 17 machines | The paper's text says 16 machines and a table caption says 17; instances are messages grouped within 10 seconds. | talagala-ucb99-pdf |
| Share of errors that were name-lookup and file-mount errors | not stated | 40 | percent of error instances (lower bound) | Same logs | Printed as over 40%; caused by dependence on external servers. | talagala-ucb99-pdf |
| Share of errors that were SCSI timeouts or parity errors | 49 | 87 | percent of error instances | 49% of all errors; 87% once network errors are removed | Mostly one enclosure failure and loose or faulty cables. | talagala-ucb99-pdf |
| Share of errors that came from data disks | not stated | 4 | percent of error instances (about) | Same logs | Printed as around 4% overall, with about 3% recovered errors; the paper's table lists fewer. | talagala-ucb99-pdf |
| Restarts caused by data disk or IDE disk errors | not stated | 0 | restarts of 73 | 16 nodes, 6 months | The largest cause of restarts was external network failures. | talagala-ucb99-pdf |
| Disks as a share of hardware replacements | 20 | 50 | percent of replacements | HPC1 30%, COM2 50%, COM1 nearly 20% | Counts of replacements, not a rate per disk. | schroeder-fast07-pdf |
| Disk mean time to failure | 10 | 50 | years | Tens of Google storage cells | Counts software and hardware causes. | ford-osdi10-pdf |
| Node mean time to failure | not stated | 4.3 | months | Same cells | Most node failures are transient. | ford-osdi10-pdf |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'over' or 'around' it is the stated figure.
- unit
- Unit of the value columns: disks, months, components, percent, restarts, or years.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Talagala and Patterson (1999), Schroeder and Gibson (FAST 2007), and Ford et al. (OSDI 2010), with credit. Not a Creative Commons license. The papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- An Analysis of Error Behavior in a Large Storage System, Talagala and Patterson, UC Berkeley technical report UCB/CSD-99-1042, February 1999 (author PDF, 18 pages)
- Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4)
- Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF, section 3)