Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · disk error prediction · ATC 2018

Azure disk-error prediction caught 36.5% of faulty disks at a 0.1% false-positive rate, and random cross-validation said 91.6%

Xu et al., USENIX ATC 2018, built a disk-error predictor for Azure. At a 0.1% false-positive rate it caught 36.50% of faulty disks on one test set (29.67% to 41.09% across three), against 15.51% to 30.51% for SVM and random forest baselines on SMART data; a random split gave 91.64% on the same data. Rows from Backblaze (2016) show the SMART gap in another fleet.

Download CSV15 rows · 7 columns · CSV

File /data/datasets/cloud-disk-error-prediction.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Xu et al. (USENIX ATC 2018) and two cross-check rows on SMART warnings from Backblaze (2016). It does not read chart values. TPR is the share of faulty disks flagged. The false-positive rate is held at 0.1% of healthy disks. Results are one cloud system's own, tested on the month after training.

Limits

Table

15 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-10.
MeasureValue lowValue highUnitScopeNoteCitation
healthy to faulty disks in each test setnot stated3faulty disks per 10000 healthyThree test sets from November 2017Printed as around 10000 to 3.xu-atc18-pdf
disks that become faulty per daynot stated300disks per 1000000The studied cloud systemPrinted as about 300 out of 1000000.xu-atc18-pdf
CDEF true positive rate at 0.1% false positives Dataset 1not stated36.50percent of faulty disksCost-sensitive ranking modelxu-atc18-pdf
CDEF true positive rate at 0.1% false positives Dataset 2not stated41.09percent of faulty disksSame modelxu-atc18-pdf
CDEF true positive rate at 0.1% false positives Dataset 3not stated29.67percent of faulty disksSame modelxu-atc18-pdf
random forest true positive rate at 0.1% false positives Dataset 1not stated30.51percent of faulty disksSMART features onlyDatasets 2 and 3 were 34.11 and 18.81.xu-atc18-pdf
SVM true positive rate at 0.1% false positives Dataset 1not stated15.51percent of faulty disksSMART features onlyDatasets 2 and 3 were 21.71 and 7.20.xu-atc18-pdf
AUC on Dataset 1 for CDEFnot stated0.93AUCA random guess is 0.5Random forest 0.85 and SVM 0.53.xu-atc18-pdf
average true positive rate with SMART features alonenot stated27.6percent of faulty disksAt 0.1% false positives and three datasetsxu-atc18-pdf
average true positive rate with SMART plus system-level signalsnot stated30.3percent of faulty disksSame settingxu-atc18-pdf
average true positive rate with feature selection addednot stated35.8percent of faulty disksSame setting and this is CDEFxu-atc18-pdf
true positive rate on Dataset 1 with cross-validationnot stated91.64percent of faulty disksAt 0.1% false positivesThe time split gave 36.50 on the same data.xu-atc18-pdf
average VM downtime saved after deploymentnot stated63000minutes per monthAzurePrinted as around 63k. 99.999% availability allows 26 seconds per month.xu-atc18-pdf
Backblaze failed drives with one or more of five SMART counts above zeronot stated76.7percent of failed drives67814 drives in 201623.3% showed none.backblaze-smart-2016
Backblaze operational drives with one or more of five SMART counts above zeronot stated4.2percent of operational drivesSame fleetbackblaze-smart-2016

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
unit (string)
Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Xu et al. (USENIX ATC 2018) and a Backblaze post of 6 October 2016, with credit. Not a Creative Commons license. The USENIX proceedings state that rights to individual papers remain with the author or the author's employer. This file is not a copy of the paper and not the data.

Sources

  1. Improving Service Availability of Cloud Systems by Predicting Disk Error, Xu et al., USENIX ATC 2018 (USENIX PDF, pages 481 to 494) (accessed 2026-10-10)
  2. What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 (accessed 2026-10-10)