Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · SSD failures · SYSTOR 2016

38% of failed datacenter SSDs showed none of the four SMART symptoms; the ones that did failed 3 to 20 times as often

Narayanan et al., SYSTOR 2016, studied over half a million SSDs in one operator's datacenters over nearly 3 years. 62% of failed devices had shown at least one of four SMART symptoms and 38% none; devices with a symptom failed 3 to 20 times as often. A classifier reached 87% precision and 71% recall with 5-fold cross-validation. Cross-check rows come from NetApp (FAST 2020) and Backblaze (2016).

Download CSV14 rows · 7 columns · CSV

File /data/datasets/ssd-failures-datacenter-symptoms.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.

Limits

Table

14 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-10.
MeasureValue lowValue highUnitScopeNoteCitation
SSDs studied500000devices5 model groups and one operatorPrinted as over half a million.narayanan-systor16-pdf
vendor-quoted annual failure rate for consumer models0.610.73percent per yearSpecification range shown in the AFR chartnarayanan-systor16-pdf
observed AFR above specification for some modelsnot stated70percent above the quoted rateModels 1-B and 1-C exceeded the rangePrinted as as much as 70.narayanan-systor16-pdf
replacement after an SSD failure ticketnot stated79percent of ticketsDatacenter tickets in the studyHard disk tickets were 11%.narayanan-systor16-pdf
AFR increase when a symptom is present320timesFour SMART symptom categoriesData errors up to 20 times.narayanan-systor16-pdf
AFR increase with reallocations or SATA downshiftnot stated4timesCompared with devices without the symptomnarayanan-systor16-pdf
AFR increase with program or erase failuresnot stated2.75timesCompared with devices without the symptomnarayanan-systor16-pdf
failed devices with at least one of the four symptomsnot stated62percent of failed devicesFail-stop failuresPrinted as around 62. The other 38% showed none.narayanan-systor16-pdf
healthy devices with data errorsnot stated1percent of healthy devicesSame fleetAbout 12% of failed devices had them.narayanan-systor16-pdf
AFR increase with average writes per day24timesOlder models 1-A and 1-B and 1-CNot seen for model 1-D.narayanan-systor16-pdf
classifier precisionnot stated87percentFailed versus healthy devices5-fold cross-validation with SMOTE over-sampling.narayanan-systor16-pdf
classifier recallnot stated71percentSame classifiernarayanan-systor16-pdf
average annual replacement rate of enterprise SSDsnot stated0.22percent per yearNetApp and about 1.4 million SSDsModels ranged from 0.07 to nearly 1.2.maneas-fast20-pdf
Backblaze failed drives with one or more of five SMART counts above zeronot stated76.7percent of failed drives67814 drives in 201623.3% showed none.backblaze-smart-2016

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
unit (string)
Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Narayanan et al. (SYSTOR 2016), Maneas et al. (FAST 2020), and a Backblaze post of 6 October 2016, with credit. Not a Creative Commons license. The ACM and USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.

Sources

  1. SSD Failures in Datacenters: What? When? and Why?, Narayanan et al., SYSTOR 2016 (author PDF, ACM article 7, pages 1 to 11) (accessed 2026-10-10)
  2. A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, pages 137 to 149) (accessed 2026-10-10)
  3. What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 (accessed 2026-10-10)