Dataset · SSD failures · SYSTOR 2016
38% of failed datacenter SSDs showed none of the four SMART symptoms; the ones that did failed 3 to 20 times as often
Narayanan et al., SYSTOR 2016, studied over half a million SSDs in one operator's datacenters over nearly 3 years. 62% of failed devices had shown at least one of four SMART symptoms and 38% none; devices with a symptom failed 3 to 20 times as often. A classifier reached 87% precision and 71% recall with 5-fold cross-validation. Cross-check rows come from NetApp (FAST 2020) and Backblaze (2016).
Download CSV14 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.
Limits
- One operator's fleet and workloads; model names are anonymized.
- Failure means a fail-stop that takes the server down; about 80% ended in replacement.
- Per-model AFR values are in charts, so only printed text figures are used.
- The classifier was scored with random 5-fold cross-validation, not a split by time.
- The cross-check rows are other fleets with other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| SSDs studied | 500000 | devices | 5 model groups and one operator | Printed as over half a million. | narayanan-systor16-pdf | |
| vendor-quoted annual failure rate for consumer models | 0.61 | 0.73 | percent per year | Specification range shown in the AFR chart | narayanan-systor16-pdf | |
| observed AFR above specification for some models | not stated | 70 | percent above the quoted rate | Models 1-B and 1-C exceeded the range | Printed as as much as 70. | narayanan-systor16-pdf |
| replacement after an SSD failure ticket | not stated | 79 | percent of tickets | Datacenter tickets in the study | Hard disk tickets were 11%. | narayanan-systor16-pdf |
| AFR increase when a symptom is present | 3 | 20 | times | Four SMART symptom categories | Data errors up to 20 times. | narayanan-systor16-pdf |
| AFR increase with reallocations or SATA downshift | not stated | 4 | times | Compared with devices without the symptom | narayanan-systor16-pdf | |
| AFR increase with program or erase failures | not stated | 2.75 | times | Compared with devices without the symptom | narayanan-systor16-pdf | |
| failed devices with at least one of the four symptoms | not stated | 62 | percent of failed devices | Fail-stop failures | Printed as around 62. The other 38% showed none. | narayanan-systor16-pdf |
| healthy devices with data errors | not stated | 1 | percent of healthy devices | Same fleet | About 12% of failed devices had them. | narayanan-systor16-pdf |
| AFR increase with average writes per day | 2 | 4 | times | Older models 1-A and 1-B and 1-C | Not seen for model 1-D. | narayanan-systor16-pdf |
| classifier precision | not stated | 87 | percent | Failed versus healthy devices | 5-fold cross-validation with SMOTE over-sampling. | narayanan-systor16-pdf |
| classifier recall | not stated | 71 | percent | Same classifier | narayanan-systor16-pdf | |
| average annual replacement rate of enterprise SSDs | not stated | 0.22 | percent per year | NetApp and about 1.4 million SSDs | Models ranged from 0.07 to nearly 1.2. | maneas-fast20-pdf |
| Backblaze failed drives with one or more of five SMART counts above zero | not stated | 76.7 | percent of failed drives | 67814 drives in 2016 | 23.3% showed none. | backblaze-smart-2016 |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
- unit
- Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Narayanan et al. (SYSTOR 2016), Maneas et al. (FAST 2020), and a Backblaze post of 6 October 2016, with credit. Not a Creative Commons license. The ACM and USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- SSD Failures in Datacenters: What? When? and Why?, Narayanan et al., SYSTOR 2016 (author PDF, ACM article 7, pages 1 to 11)
- A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, pages 137 to 149)
- What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016