Dataset · disk failure prediction · FAST 2020
Adding performance and location data lifted 10-day disk failure prediction to 0.95 MCC across 380,000 disks
Lu et al., FAST 2020, studied 380,000 hard disks in 64 sites of one operator over about 70 days. With SMART, performance, and location data together, a CNN-LSTM scored 0.95 MCC for a 10-day prediction horizon, against 0.77 for the next best method, a random forest. Rows from Backblaze (2016) and Google (FAST 2007) show the SMART gap in other fleets.
Download CSV13 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Lu et al. (FAST 2020) and two cross-check rows on SMART warnings from Backblaze (2016) and Google (FAST 2007). It does not read chart values. MCC is the Matthews correlation coefficient, from -1 to 1. The scores are one operator's own results for a 10-day horizon.
Limits
- Scores are the paper's own, from one operator, about 70 days, and a 10-day horizon.
- The number of failed disks is not stated in the text read.
- The paper prints 2.6 million device hours, which does not match 380,000 disks over about 70 days.
- Portability across sites held for the CNN-LSTM but not always for RF and GBDT.
- The Backblaze and Google rows measure SMART warnings, not model performance.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| hard disks in the study | not stated | 380000 | disks | 64 sites and 10000 server racks of one operator | The operator houses more than two million disks. | lu-fast20-pdf |
| observation window | not stated | 70 | days | SMART collected once a day | Printed as roughly 70. | lu-fast20-pdf |
| prediction horizon | not stated | 10 | days | All main results | Chosen so operators have time to act. | lu-fast20-pdf |
| best model MCC with SMART and performance and location data | not stated | 0.95 | MCC score | CNN-LSTM and SPL feature group | Printed as up to 0.95. | lu-fast20-pdf |
| best model F-measure | not stated | 0.95 | F-measure | CNN-LSTM and 10-day horizon | Abstract says on average and introduction says up to. | lu-fast20-pdf |
| next best method MCC with the same data | not stated | 0.77 | MCC score | Random forest and SPL feature group | lu-fast20-pdf | |
| gain from adding location information | not stated | 10 | percent of MCC | CNN-LSTM | Printed as less than 10 so 10 is an upper bound. | lu-fast20-pdf |
| MCC on an unseen site after training on 62 sites | 0.90 | MCC score | CNN-LSTM and SPL group and site A | Printed as above 0.90 so 0.90 is a lower bound. | lu-fast20-pdf | |
| drop for RF and GBDT on an unseen site | 15 | percent in some cases | Same test | Printed as more than 15. | lu-fast20-pdf | |
| training time of the CNN-LSTM | not stated | 4 | hours per training run | Same study | Printed as up to four hours. | lu-fast20-pdf |
| Backblaze failed drives with one or more of five SMART counts above zero | not stated | 76.7 | percent of failed drives | 67814 drives in 2016 | 23.3% of failed drives showed none. | backblaze-smart-2016 |
| Backblaze operational drives with one or more of five SMART counts above zero | not stated | 4.2 | percent of operational drives | Same fleet | backblaze-smart-2016 | |
| Google failed drives with no count on four strong SMART signals | 56 | percent of failed drives | More than 100000 ATA disks from December 2005 to August 2006 | Printed as over 56. | pinheiro-usenix-html |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
- unit
- Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Lu et al. (FAST 2020), Backblaze (6 October 2016), and Pinheiro et al. (FAST 2007), with credit. Not a Creative Commons license. The USENIX proceedings state that rights to individual papers remain with the author or the author's employer. This file is not a copy of the papers and not the data set.
Sources
- Making Disk Failure Predictions SMARTer!, Lu, Luo, Patel, Yao, Tiwari, and Shi, FAST 2020 (USENIX PDF, pages 151 to 167)
- What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016
- Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings)