Dataset · SSD failures · SYSTOR 2016
A 2016 SMART predictor reached 98% recall on one Seagate model but 81% on a Hitachi model, scored with random splits on Backblaze data
Botezatu et al., KDD 2016, predicted disk replacement from SMART data of 50,984 Backblaze disks: 98% recall on one Seagate model and 81% on one Hitachi model, scored with 100 random 80/20 splits. Cross-check rows come from Azure (ATC 2018), Backblaze (2016), and a data center operator (FAST 2020).
Download CSV15 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.
Limits
- Two disk families from one public fleet, 2013 to 2015; a model is needed per manufacturer.
- Scores come from random 80/20 splits of a downsampled set, not from a split by time.
- The paper does not say the test set keeps the true 2.5% to 3% share of replaced disks.
- The Backblaze post is the same fleet's own analysis, not an independent dataset.
- The cross-check rows use other data and other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| Disks in the Backblaze dataset used | not stated | 50984 | disks | Backblaze public daily SMART logs | Only a Hitachi and a Seagate family were kept; other makers had too few samples or too few SMART values. | botezatu-kdd16-pdf |
| Months of data kept | not stated | 17 | months | April 2013 to June 2015 collection; first months dropped | Over 70% of SMART values were not collected before January 2014. | botezatu-kdd16-pdf |
| Replaced share of the two main models | 2.5 | 3 | percent of disks | Seagate ST4000DM000 (SgtA) and Hitachi HDS722020ALA330 (HitA) | The healthy class was downsampled to 1,000 (SgtA) and 500 (HitA) before training. | botezatu-kdd16-pdf |
| Error on the best Seagate model | 1 | 2 | percent over 100 runs | SgtA, regularized greedy forest | Printed as 98% accuracy and 1-2% error. | botezatu-kdd16-pdf |
| Recall for replaced disks, Seagate model | not stated | 98 | percent | SgtA, 100 random 80/20 splits | Median over runs. | botezatu-kdd16-pdf |
| Recall for replaced disks, Hitachi model | not stated | 81 | percent | HitA, 100 random 80/20 splits | Precision, recall, and F-score for Hitachi are 14 to 19% lower than for Seagate. | botezatu-kdd16-pdf |
| Recall of a simple decision tree on a small set of SMART attributes | not stated | 53 | percent | SgtA; 44% for HitA | The authors' baseline with a commonly used subset of attributes. | botezatu-kdd16-pdf |
| Replaced Seagate disks predicted 10 days ahead | not stated | 92 | percent | SgtA; 97% at 3 days | From a snapshot taken 10 days before replacement. | botezatu-kdd16-pdf |
| Replaced Seagate disks predicted 30 days ahead | not stated | 73 | percent | SgtA; HitA 75% at 30 days | From a snapshot taken 30 days before replacement. | botezatu-kdd16-pdf |
| Evaluation repeats with random training and test splits | not stated | 100 | splits of 80% training and 20% test | Same models | Not split by time. | botezatu-kdd16-pdf |
| True positive rate with random cross-validation at a 0.1% false positive rate | not stated | 91.64 | percent | Microsoft Azure, Dataset 1 | Xu et al. cite Botezatu et al. among the studies that use cross-validation. | xu-atc18-pdf |
| True positive rate with online prediction at the same false positive rate | not stated | 36.5 | percent | Same dataset and model, training before testing in time | The drop is the authors' evidence that random splits are optimistic. | xu-atc18-pdf |
| Failed drives with at least one of five SMART counts above zero | not stated | 76.7 | percent of failed drives | Backblaze, 67,814 drives, 2016 | 23.3% showed no warning; the post is the fleet's own analysis, not an independent dataset. | backblaze-smart-2016 |
| Operational drives with at least one of the five SMART counts above zero | not stated | 4.2 | percent of operational drives | Same fleet | Counts above zero alone are a weak alarm. | backblaze-smart-2016 |
| Matthews correlation coefficient with 5-fold cross-validation | not stated | 0.95 | correlation coefficient | disks of a leading data center operator, 10-day lead time | Also a random partition, so the same caution applies. | lu-fast20-pdf |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range.
- unit
- Unit of the value columns: disks, months, percent, splits, or correlation coefficient.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Botezatu et al. (KDD 2016), Xu et al. (ATC 2018), a Backblaze post of 6 October 2016, and Lu et al. (FAST 2020), with credit. Not a Creative Commons license. The papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- Predicting Disk Replacement towards Reliable Data Centers, Botezatu, Giurgiu, Bogojeska, and Wiesmann, ACM SIGKDD 2016, pages 39 to 48 (KDD PDF)
- Improving Service Availability of Cloud Systems by Predicting Disk Error, Xu et al., USENIX ATC 2018 (USENIX PDF, section 3 on online prediction)
- What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016
- Making Disk Failure Predictions SMARTer!, Lu et al., USENIX FAST 2020 (USENIX PDF, section 4)