Dataset · SSD failures · SYSTOR 2016
In Alibaba's SSD fleet, 12.9% of failures were in the same node and 18.3% in the same rack within 30 minutes of another failure
Han et al., FAST 2021, studied nearly one million Alibaba SSDs over two years. 12.9% of failures were intra-node and 18.3% intra-rack within 30 minutes, and the strongest SMART attribute had a rank correlation of only 0.23 with them. Cross-check rows come from NetApp (FAST 2020) and Google (OSDI 2010).
Download CSV16 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.
Limits
- One operator's fleet and ticketing rules.
- The 12.9% and 18.3% shares depend on the 30 minute threshold.
- 88.6% of nodes with two or more SSDs hold one drive model, so some clustering may be a model effect.
- The redundancy results come from a trace-driven simulation.
- The cross-check rows are other fleets with other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| SSDs studied | not stated | 1000000 | devices | 11 drive models from 3 vendors, Alibaba data centers | Printed as nearly one million. | han-fast21-pdf |
| Nodes holding the SSDs | not stated | 200000 | nodes | Same fleet | Printed as 200 K nodes; 88.6% of nodes with at least two SSDs hold one drive model. | han-fast21-pdf |
| Racks holding the SSDs | not stated | 30000 | racks | Same fleet | Printed as 30 K racks. | han-fast21-pdf |
| Length of the record | not stated | 2 | years | January 2018 to December 2019 | Daily SMART logs, trouble tickets, locations, applications. | han-fast21-pdf |
| Failed SSDs (trouble tickets) | not stated | 19000 | drives | Same fleet | Printed as about 19 K; whole-drive and partial-drive failures, each checked by an administrator. | han-fast21-pdf |
| Annual failure rate of all SSDs | not stated | 1.16 | percent per year | Same fleet | Computed from the tickets and drive-days. | han-fast21-pdf |
| Failures that are intra-node failures | not stated | 12.9 | percent of failures | Failures in one node within 30 minutes of each other | Default 30 minute threshold. | han-fast21-pdf |
| Failures that are intra-rack failures | not stated | 18.3 | percent of failures | Failures in one rack within 30 minutes of each other | Default 30 minute threshold. | han-fast21-pdf |
| Chance of one more failure in an intra-node failure group | 26.3 | 64.3 | percent | Group size 2 to 11 failures | If failures were independent this would be close to the annual failure rate of 1.16%. | han-fast21-pdf |
| Intra-node failures with an interval of one minute or less | not stated | 10.0 | percent of intra-node failures | Same fleet | The one month threshold gives 29.2%. | han-fast21-pdf |
| Intra-rack failures with an interval of one month or less | not stated | 63.0 | percent of intra-rack failures | Same fleet | The one minute threshold gives 14.4%. | han-fast21-pdf |
| Highest rank correlation of a SMART attribute with correlated failures | not stated | 0.23 | Spearman correlation coefficient | SMART attribute S187 (reported uncorrectable errors) | Same value for intra-node and intra-rack failures; the paper calls this limited. | han-fast21-pdf |
| Intra-node failure share by application | 2.1 | 33.6 | percent of an application's SSD failures | Eight largest applications | Intra-rack share ranges from 2.8% to 40.5%. | han-fast21-pdf |
| Chance of a drive replacement in a random week | not stated | 0.0504 | percent | NetApp RAID groups, about 1.4 million SSDs | Baseline for the next row. | maneas-fast20-pdf |
| Chance of a replacement within a week of a previous one in the same RAID group | not stated | 9.39 | percent | Same fleet | More than 180 times the baseline; 52% of consecutive replacements fall within a week. | maneas-fast20-pdf |
| Node failures that are part of a burst of at least 2 nodes | not stated | 37 | percent of failures | Tens of Google storage cells, 120 second window | Printed as 37%; node unavailability, not SSD failures. | ford-osdi10-pdf |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'nearly' or 'about' it is the stated figure.
- unit
- Unit of the value columns: devices, nodes, racks, years, percent, or correlation coefficient.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Han et al. (FAST 2021), Maneas et al. (FAST 2020), and Ford et al. (OSDI 2010), with credit. Not a Creative Commons license. The USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- An In-Depth Study of Correlated Failures in Production SSD-Based Data Centers, Han, Lee, Xu, Liu, He, and Liu, USENIX FAST 2021, pages 417 to 429 (USENIX PDF)
- A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements)
- Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF, section 4 on correlated failures)