Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · SSD failures · SYSTOR 2016

In Alibaba's SSD fleet, 12.9% of failures were in the same node and 18.3% in the same rack within 30 minutes of another failure

Han et al., FAST 2021, studied nearly one million Alibaba SSDs over two years. 12.9% of failures were intra-node and 18.3% intra-rack within 30 minutes, and the strongest SMART attribute had a rank correlation of only 0.23 with them. Cross-check rows come from NetApp (FAST 2020) and Google (OSDI 2010).

Download CSV16 rows · 7 columns · CSV

File /data/datasets/ssd-correlated-failures-nodes-racks.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.

Limits

Table

16 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-11.
MeasureValue lowValue highUnitScopeNoteCitation
SSDs studiednot stated1000000devices11 drive models from 3 vendors, Alibaba data centersPrinted as nearly one million.han-fast21-pdf
Nodes holding the SSDsnot stated200000nodesSame fleetPrinted as 200 K nodes; 88.6% of nodes with at least two SSDs hold one drive model.han-fast21-pdf
Racks holding the SSDsnot stated30000racksSame fleetPrinted as 30 K racks.han-fast21-pdf
Length of the recordnot stated2yearsJanuary 2018 to December 2019Daily SMART logs, trouble tickets, locations, applications.han-fast21-pdf
Failed SSDs (trouble tickets)not stated19000drivesSame fleetPrinted as about 19 K; whole-drive and partial-drive failures, each checked by an administrator.han-fast21-pdf
Annual failure rate of all SSDsnot stated1.16percent per yearSame fleetComputed from the tickets and drive-days.han-fast21-pdf
Failures that are intra-node failuresnot stated12.9percent of failuresFailures in one node within 30 minutes of each otherDefault 30 minute threshold.han-fast21-pdf
Failures that are intra-rack failuresnot stated18.3percent of failuresFailures in one rack within 30 minutes of each otherDefault 30 minute threshold.han-fast21-pdf
Chance of one more failure in an intra-node failure group26.364.3percentGroup size 2 to 11 failuresIf failures were independent this would be close to the annual failure rate of 1.16%.han-fast21-pdf
Intra-node failures with an interval of one minute or lessnot stated10.0percent of intra-node failuresSame fleetThe one month threshold gives 29.2%.han-fast21-pdf
Intra-rack failures with an interval of one month or lessnot stated63.0percent of intra-rack failuresSame fleetThe one minute threshold gives 14.4%.han-fast21-pdf
Highest rank correlation of a SMART attribute with correlated failuresnot stated0.23Spearman correlation coefficientSMART attribute S187 (reported uncorrectable errors)Same value for intra-node and intra-rack failures; the paper calls this limited.han-fast21-pdf
Intra-node failure share by application2.133.6percent of an application's SSD failuresEight largest applicationsIntra-rack share ranges from 2.8% to 40.5%.han-fast21-pdf
Chance of a drive replacement in a random weeknot stated0.0504percentNetApp RAID groups, about 1.4 million SSDsBaseline for the next row.maneas-fast20-pdf
Chance of a replacement within a week of a previous one in the same RAID groupnot stated9.39percentSame fleetMore than 180 times the baseline; 52% of consecutive replacements fall within a week.maneas-fast20-pdf
Node failures that are part of a burst of at least 2 nodesnot stated37percent of failuresTens of Google storage cells, 120 second windowPrinted as 37%; node unavailability, not SSD failures.ford-osdi10-pdf

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'nearly' or 'about' it is the stated figure.
unit (string)
Unit of the value columns: devices, nodes, racks, years, percent, or correlation coefficient.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Han et al. (FAST 2021), Maneas et al. (FAST 2020), and Ford et al. (OSDI 2010), with credit. Not a Creative Commons license. The USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.

Sources

  1. An In-Depth Study of Correlated Failures in Production SSD-Based Data Centers, Han, Lee, Xu, Liu, He, and Liu, USENIX FAST 2021, pages 417 to 429 (USENIX PDF) (accessed 2026-10-11)
  2. A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements) (accessed 2026-10-11)
  3. Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF, section 4 on correlated failures) (accessed 2026-10-11)