Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · SSD failures · SYSTOR 2016

80% of 315 verified fail-slow drives at Alibaba were caused by software scheduling, and 216 of them were in two clusters

Lu et al., FAST 2023 (Perseus), monitored 248K Alibaba drives for 10 months. Of 315 verified fail-slow drives, 252 were caused by software scheduling and 63 by hardware, and 216 of the 252 were in two clusters. Cross-check rows come from Gunawi et al., FAST 2018 (101 reports, 12 institutions).

Download CSV15 rows · 7 columns · CSV

File /data/datasets/fail-slow-drives-software-scheduling.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.

Limits

Table

15 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-11.
MeasureValue lowValue highUnitScopeNoteCitation
Drives under close monitoringnot stated248000drivesAlibaba cloud storage, one 10 month windowPrinted as 248K drives.lu-fast23-pdf
Length of the monitoring windownot stated10monthsSame drivesPrinted as 10-month close monitoring.lu-fast23-pdf
Fail-slow cases found by the detectornot stated304drivesSame drivesAll 304 are among the verified drives below.lu-fast23-pdf
Verified fail-slow drives in the benchmarknot stated315drives25 clusters; 237 SSDs and 78 HDDsVerified by on-site engineers or manufacturers.lu-fast23-pdf
Normal peer drives in the benchmarknot stated41000drivesSame clustersPrinted as around 41K.lu-fast23-pdf
Verified fail-slow drives caused by software schedulingnot stated252drives216 SSDs and 36 HDDs252 of 315 is 80.0% by division; 216 SSDs were in two clusters with the same logical drive IDs.lu-fast23-pdf
Verified fail-slow drives with hardware causesnot stated63drives42 HDDs and 21 SSDs63 of 315 is 20.0% by division.lu-fast23-pdf
Hardware cases with a root cause returned by the vendornot stated15drives9 HDDs and 6 SSDsDiagnosis was lengthy, so most hardware cases have no vendor root cause.lu-fast23-pdf
Cut in node-level 95th percentile write latency after isolating fail-slow drivesnot stated30.67percentMean, plus or minus 10.96Not a per-drive figure.lu-fast23-pdf
Cut in node-level 99.99th percentile write latency after isolating fail-slow drivesnot stated48.05percentMean, plus or minus 15.53Printed as 48% in the abstract.lu-fast23-pdf
Detector precision on the benchmarknot stated0.99fractionFull test dataset, 315 positivesIncludes clusters with software-caused cases.lu-fast23-pdf
Detector recall on the benchmarknot stated1.00fractionSame datasetThe benchmark is labeled from drives the detector flagged and others verified.lu-fast23-pdf
Fail-slow hardware incident reports collectednot stated101reports12 institutions, large-scale clustersA report can hold several causes, so causes total 112.gunawi-fast18-pdf
Root causes that are external factorsnot stated39percent of root causesConfiguration, environment, power, temperaturePrinted as 39%; the rest are internal firmware or device errors or unknown.gunawi-fast18-pdf
Fail-slow incidents that took months to detectnot stated17percent of incidentsSame reports; 45% have an unknown time13% were found in hours, 13% in days, 11% in weeks.gunawi-fast18-pdf

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'about' it is the stated figure.
unit (string)
Unit of the value columns: drives, months, reports, percent, or fraction.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Lu et al. (FAST 2023) and Gunawi et al. (FAST 2018), with credit. Not a Creative Commons license. The USENIX papers remain with their authors or publishers. This file is not a copy of the papers and not the data.

Sources

  1. Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems, Lu, Xu, Zhang, Zhu, Zhu, Wang, Zhu, Xue, Shu, Li, and Wu, USENIX FAST 2023 (USENIX PDF, sections 5 and 6) (accessed 2026-10-11)
  2. Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems, Gunawi et al., USENIX FAST 2018 (USENIX PDF, sections 2 to 3) (accessed 2026-10-11)