Dataset · disk error prediction · ATC 2018
Azure disk-error prediction caught 36.5% of faulty disks at a 0.1% false-positive rate, and random cross-validation said 91.6%
Xu et al., USENIX ATC 2018, built a disk-error predictor for Azure. At a 0.1% false-positive rate it caught 36.50% of faulty disks on one test set (29.67% to 41.09% across three), against 15.51% to 30.51% for SVM and random forest baselines on SMART data; a random split gave 91.64% on the same data. Rows from Backblaze (2016) show the SMART gap in another fleet.
Download CSV15 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Xu et al. (USENIX ATC 2018) and two cross-check rows on SMART warnings from Backblaze (2016). It does not read chart values. TPR is the share of faulty disks flagged. The false-positive rate is held at 0.1% of healthy disks. Results are one cloud system's own, tested on the month after training.
Limits
- One cloud system of one company, one month of training and one month of testing.
- Labels come from engineers' root-cause analyses and some disks return from repair.
- Only TPR, FPR, and AUC are reported; precision is not.
- The 63000 minutes saved is the authors' own average after deployment.
- The Backblaze rows measure SMART warnings in another fleet, not this model.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| healthy to faulty disks in each test set | not stated | 3 | faulty disks per 10000 healthy | Three test sets from November 2017 | Printed as around 10000 to 3. | xu-atc18-pdf |
| disks that become faulty per day | not stated | 300 | disks per 1000000 | The studied cloud system | Printed as about 300 out of 1000000. | xu-atc18-pdf |
| CDEF true positive rate at 0.1% false positives Dataset 1 | not stated | 36.50 | percent of faulty disks | Cost-sensitive ranking model | xu-atc18-pdf | |
| CDEF true positive rate at 0.1% false positives Dataset 2 | not stated | 41.09 | percent of faulty disks | Same model | xu-atc18-pdf | |
| CDEF true positive rate at 0.1% false positives Dataset 3 | not stated | 29.67 | percent of faulty disks | Same model | xu-atc18-pdf | |
| random forest true positive rate at 0.1% false positives Dataset 1 | not stated | 30.51 | percent of faulty disks | SMART features only | Datasets 2 and 3 were 34.11 and 18.81. | xu-atc18-pdf |
| SVM true positive rate at 0.1% false positives Dataset 1 | not stated | 15.51 | percent of faulty disks | SMART features only | Datasets 2 and 3 were 21.71 and 7.20. | xu-atc18-pdf |
| AUC on Dataset 1 for CDEF | not stated | 0.93 | AUC | A random guess is 0.5 | Random forest 0.85 and SVM 0.53. | xu-atc18-pdf |
| average true positive rate with SMART features alone | not stated | 27.6 | percent of faulty disks | At 0.1% false positives and three datasets | xu-atc18-pdf | |
| average true positive rate with SMART plus system-level signals | not stated | 30.3 | percent of faulty disks | Same setting | xu-atc18-pdf | |
| average true positive rate with feature selection added | not stated | 35.8 | percent of faulty disks | Same setting and this is CDEF | xu-atc18-pdf | |
| true positive rate on Dataset 1 with cross-validation | not stated | 91.64 | percent of faulty disks | At 0.1% false positives | The time split gave 36.50 on the same data. | xu-atc18-pdf |
| average VM downtime saved after deployment | not stated | 63000 | minutes per month | Azure | Printed as around 63k. 99.999% availability allows 26 seconds per month. | xu-atc18-pdf |
| Backblaze failed drives with one or more of five SMART counts above zero | not stated | 76.7 | percent of failed drives | 67814 drives in 2016 | 23.3% showed none. | backblaze-smart-2016 |
| Backblaze operational drives with one or more of five SMART counts above zero | not stated | 4.2 | percent of operational drives | Same fleet | backblaze-smart-2016 |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
- unit
- Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Xu et al. (USENIX ATC 2018) and a Backblaze post of 6 October 2016, with credit. Not a Creative Commons license. The USENIX proceedings state that rights to individual papers remain with the author or the author's employer. This file is not a copy of the paper and not the data.