Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · reallocated sectors · FAST 2015

Replacing disks at 200 reallocated sectors removed 88% of triple-disk RAID failures at one vendor

Ma et al., FAST 2015, analysed about 1 million SATA disks in 6 models in one vendor's backup systems (logs of 21 to 60 months from June 2008). Replacing a disk once its reallocated-sector count passed 200 avoided about 88% of triple-disk RAID failures, which were 80% of RAID failures, or about 70% of all disk-related incidents. In simulation, thresholds below 200 caught 52% to 70% of failures within 60 days with 0.8% to 4.5% false positives. Cross-check rows come from Google (FAST 2007) and Backblaze (2016).

Download CSV17 rows · 7 columns · CSV

File /data/datasets/raidshield-reallocated-sectors.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Ma et al. (FAST 2015) and two cross-check sets from Google (FAST 2007) and Backblaze (2016). It does not read chart values. Reallocated sectors (RS) are sectors a disk remapped to a spare after a write error. A triple failure is three simultaneous whole-disk or sector failures in one RAID-6 group. The 88% is a vendor's own before and after comparison.

Limits

Table

17 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-10.
MeasureValue lowValue highUnitScopeNoteCitation
triple failures among RAID failures before the rulenot stated80percent of RAID failuresVendor backup systemsAnother 5% were other hardware faults and 15% user errors and unknown causes.ma-fast15-pdf
triple failures avoided after replacing disks at 200 reallocated sectorsnot stated88percent of triple-failure incidentsDeployed on models A-1 and A-2 and B-1Printed as about 88. Normalized to the pre-deployment monthly average.ma-fast15-pdf
all disk-related incidents avoidednot stated70percent of incidentsSame deploymentPrinted as about 70.ma-fast15-pdf
deployed replacement thresholdnot stated200reallocated sectorsMedian time to failure fell below 3 days beyond this countChosen because replacing a disk can take up to 3 days.ma-fast15-pdf
impending failures caught with thresholds below 2005270percent of failures within 60 daysSimulation on 100000 A-2 disksThe introduction prints up to 65% instead.ma-fast15-pdf
false positives with thresholds below 2000.84.5percent of working disksSimulation on 100000 A-2 disksThe introduction prints up to 2.5% instead.ma-fast15-pdf
failed A-1 disks found in their fourth yearnot stated63percent of failed disksModel A-1Disks fail at similar ages.ma-fast15-pdf
failed A-2 disks found in their fourth yearnot stated66percent of failed disksModel A-2Disks fail at similar ages.ma-fast15-pdf
failed B-1 disks found in their fourth yearnot stated64percent of failed disksModel B-1Disks fail at similar ages.ma-fast15-pdf
growth of sector error counts in the second year25300percent increaseDisks with at least one sector error; C-2 lowest and A-2 highestPrinted as about 300 for A-2.ma-fast15-pdf
triple failures left after the rulenot stated12percent of triple failuresSudden failures or several mildly unreliable disksNot covered by the single-disk rule.ma-fast15-pdf
vulnerable RAID-6 groups flagged by ARMORnot stated80percent of vulnerable groupsSimulation on 5000 healthy and 500 failed groupsNot deployed.ma-fast15-pdf
total triple-failure coverage with ARMOR addednot stated98percent of triple failuresSimulation onlyNot deployed.ma-fast15-pdf
Google drives with a reallocation count above zeronot stated9percent of drivesMore than 100000 ATA disks from December 2005 to August 2006Printed as about 9.pinheiro-usenix-html
Google failure likelihood within 60 days after the first reallocationnot stated14times the rate for drives with noneSame populationPrinted as over 14.pinheiro-usenix-html
Backblaze failed drives with one or more of five SMART counts above zeronot stated76.7percent of failed drives67814 drives in 201623.3% of failed drives showed none.backblaze-smart-2016
Backblaze operational drives with one or more of five SMART counts above zeronot stated4.2percent of operational drives67814 drives in 2016Not a false-positive rate for a replacement rule.backblaze-smart-2016

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
unit (string)
Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Ma et al. (FAST 2015), Pinheiro et al. (FAST 2007), and a Backblaze post of 6 October 2016, with credit. Not a Creative Commons license. The USENIX proceedings state that rights to individual papers remain with the author or the author's employer. This file is not a copy of the papers and not the underlying logs.

Sources

  1. RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures, Ma et al., FAST 2015 (USENIX PDF) (accessed 2026-10-10)
  2. Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) (accessed 2026-10-09)
  3. What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 (accessed 2026-10-10)