Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · SSD failures · SYSTOR 2016

In an 18-month Berkeley study of 368 disks, 7 data disks failed against 13 enclosure backplane failures; disks were 90% of parts but about 4% of errors

Talagala and Patterson (UC Berkeley, 1999) logged 18 months of a prototype with 368 SCSI data disks. 7 data disks failed (1.9%), against 13 of 46 enclosures, and disks were about 90% of parts but around 4% of errors. Cross-check rows come from Schroeder and Gibson (FAST 2007) and Ford et al. (OSDI 2010).

Download CSV15 rows · 7 columns · CSV

File /data/datasets/disk-error-behavior-1990s-storage-prototype.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Narayanan et al. (SYSTOR 2016) and two cross-check rows from NetApp (FAST 2020) and Backblaze (2016). It does not read chart values. AFR is the share of devices with failures divided by device years. A failure here is a fail-stop that takes a server down.

Limits

Table

15 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-11.
MeasureValue lowValue highUnitScopeNoteCitation
SCSI data disks in the prototypenot stated368disks20 PCs, 3.2 TB, 8.4 GB disks, hardware from 1996Disks were about 90% of the components.talagala-ucb99-pdf
Length of the failure recordnot stated18monthsPrototype operationError logs cover the last 6 months.talagala-ucb99-pdf
Components replaced in the failure recordnot stated34componentsAll component typesPrinted as 34 absolute failures, nearly two a month.talagala-ucb99-pdf
SCSI data disks replacednot stated1.9percent of 368 disks in 18 months7 of 368 disksLowest percentage of any component that failed.talagala-ucb99-pdf
Disk enclosure backplane failuresnot stated28.3percent of 46 enclosures in 18 months13 of 46 enclosuresBackplane integrity failures on the SCSI bus.talagala-ucb99-pdf
IDE system disks replacednot stated25.0percent of 24 disks in 18 months6 of 24 disksMounted in ordinary PC chassis, not the cooled enclosures.talagala-ucb99-pdf
Enclosure power supplies replacednot stated3.26percent of 92 supplies in 18 months3 of 92 suppliesEach enclosure had two supplies, so a failure did not take it offline.talagala-ucb99-pdf
Error instances in the logsnot stated688error instances6 months of logs, 16 or 17 machinesThe paper's text says 16 machines and a table caption says 17; instances are messages grouped within 10 seconds.talagala-ucb99-pdf
Share of errors that were name-lookup and file-mount errorsnot stated40percent of error instances (lower bound)Same logsPrinted as over 40%; caused by dependence on external servers.talagala-ucb99-pdf
Share of errors that were SCSI timeouts or parity errors4987percent of error instances49% of all errors; 87% once network errors are removedMostly one enclosure failure and loose or faulty cables.talagala-ucb99-pdf
Share of errors that came from data disksnot stated4percent of error instances (about)Same logsPrinted as around 4% overall, with about 3% recovered errors; the paper's table lists fewer.talagala-ucb99-pdf
Restarts caused by data disk or IDE disk errorsnot stated0restarts of 7316 nodes, 6 monthsThe largest cause of restarts was external network failures.talagala-ucb99-pdf
Disks as a share of hardware replacements2050percent of replacementsHPC1 30%, COM2 50%, COM1 nearly 20%Counts of replacements, not a rate per disk.schroeder-fast07-pdf
Disk mean time to failure1050yearsTens of Google storage cellsCounts software and hardware causes.ford-osdi10-pdf
Node mean time to failurenot stated4.3monthsSame cellsMost node failures are transient.ford-osdi10-pdf

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'over' or 'around' it is the stated figure.
unit (string)
Unit of the value columns: disks, months, components, percent, restarts, or years.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Talagala and Patterson (1999), Schroeder and Gibson (FAST 2007), and Ford et al. (OSDI 2010), with credit. Not a Creative Commons license. The papers remain with their authors or publishers. This file is not a copy of the papers and not the data.

Sources

  1. An Analysis of Error Behavior in a Large Storage System, Talagala and Patterson, UC Berkeley technical report UCB/CSD-99-1042, February 1999 (author PDF, 18 pages) (accessed 2026-10-11)
  2. Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4) (accessed 2026-10-11)
  3. Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF, section 3) (accessed 2026-10-11)