Dataset · SSD failures · SYSTOR 2016
80% of 315 verified fail-slow drives at Alibaba were caused by software scheduling, and 216 of them were in two clusters
Lu et al., FAST 2023 (Perseus), monitored 248K Alibaba drives for 10 months. Of 315 verified fail-slow drives, 252 were caused by software scheduling and 63 by hardware, and 216 of the 252 were in two clusters. Cross-check rows come from Gunawi et al., FAST 2018 (101 reports, 12 institutions).
Download CSV15 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.
Limits
- One operator's drives and scheduling software.
- 216 of the 252 software-caused drives were in two open-channel SSD clusters, so the 80% share is concentrated.
- Only 15 of the 63 hardware cases have a vendor root cause.
- Precision and recall are on a benchmark built mostly from drives the detector flagged.
- The cross-check rows are incident reports from other institutions with other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| Drives under close monitoring | not stated | 248000 | drives | Alibaba cloud storage, one 10 month window | Printed as 248K drives. | lu-fast23-pdf |
| Length of the monitoring window | not stated | 10 | months | Same drives | Printed as 10-month close monitoring. | lu-fast23-pdf |
| Fail-slow cases found by the detector | not stated | 304 | drives | Same drives | All 304 are among the verified drives below. | lu-fast23-pdf |
| Verified fail-slow drives in the benchmark | not stated | 315 | drives | 25 clusters; 237 SSDs and 78 HDDs | Verified by on-site engineers or manufacturers. | lu-fast23-pdf |
| Normal peer drives in the benchmark | not stated | 41000 | drives | Same clusters | Printed as around 41K. | lu-fast23-pdf |
| Verified fail-slow drives caused by software scheduling | not stated | 252 | drives | 216 SSDs and 36 HDDs | 252 of 315 is 80.0% by division; 216 SSDs were in two clusters with the same logical drive IDs. | lu-fast23-pdf |
| Verified fail-slow drives with hardware causes | not stated | 63 | drives | 42 HDDs and 21 SSDs | 63 of 315 is 20.0% by division. | lu-fast23-pdf |
| Hardware cases with a root cause returned by the vendor | not stated | 15 | drives | 9 HDDs and 6 SSDs | Diagnosis was lengthy, so most hardware cases have no vendor root cause. | lu-fast23-pdf |
| Cut in node-level 95th percentile write latency after isolating fail-slow drives | not stated | 30.67 | percent | Mean, plus or minus 10.96 | Not a per-drive figure. | lu-fast23-pdf |
| Cut in node-level 99.99th percentile write latency after isolating fail-slow drives | not stated | 48.05 | percent | Mean, plus or minus 15.53 | Printed as 48% in the abstract. | lu-fast23-pdf |
| Detector precision on the benchmark | not stated | 0.99 | fraction | Full test dataset, 315 positives | Includes clusters with software-caused cases. | lu-fast23-pdf |
| Detector recall on the benchmark | not stated | 1.00 | fraction | Same dataset | The benchmark is labeled from drives the detector flagged and others verified. | lu-fast23-pdf |
| Fail-slow hardware incident reports collected | not stated | 101 | reports | 12 institutions, large-scale clusters | A report can hold several causes, so causes total 112. | gunawi-fast18-pdf |
| Root causes that are external factors | not stated | 39 | percent of root causes | Configuration, environment, power, temperature | Printed as 39%; the rest are internal firmware or device errors or unknown. | gunawi-fast18-pdf |
| Fail-slow incidents that took months to detect | not stated | 17 | percent of incidents | Same reports; 45% have an unknown time | 13% were found in hours, 13% in days, 11% in weeks. | gunawi-fast18-pdf |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'about' it is the stated figure.
- unit
- Unit of the value columns: drives, months, reports, percent, or fraction.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Lu et al. (FAST 2023) and Gunawi et al. (FAST 2018), with credit. Not a Creative Commons license. The USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems, Lu, Xu, Zhang, Zhu, Zhu, Wang, Zhu, Xue, Shu, Li, and Wu, USENIX FAST 2023 (USENIX PDF, sections 5 and 6)
- Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems, Gunawi et al., USENIX FAST 2018 (USENIX PDF, sections 2 to 3)