Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · Hardware failures · DSN 2017

Hard drives were 82% of 290,000 datacenter hardware failure tickets, and 2% of failed servers produced over 99% of the failures

Wang, Zhang, and Xu, DSN 2017, analysed over 290,000 hardware failure tickets from four years at one large Internet company. Hard drives were 81.84% of failures and 2% of failed servers produced over 99% of failures. Cross-check rows come from Schroeder and Gibson (FAST 2007) and Ford et al. (OSDI 2010).

Download CSV14 rows · 7 columns · CSV

File /data/datasets/datacenter-hardware-failure-tickets.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Wang, Zhang, and Xu (DSN 2017) and cross-check rows from Schroeder and Gibson (FAST 2007) and Ford et al. (OSDI 2010). It does not read chart values.

Limits

Table

14 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-11.
MeasureValue lowValue highUnitScopeNoteCitation
Hardware failure tickets analysednot stated290000ticketsOne large Internet company, dozens of datacenters, hundreds of thousands of serversPrinted as over 290,000 reports.wang-dsn17-pdf
Length of the recordnot stated4yearsSame companyPrinted as the past four years; the batch analysis covers 1,411 days.wang-dsn17-pdf
Hard drives as a share of all failuresnot stated81.84percent of failuresFailure tickets excluding false alarmsText says about 82%; 10.20% are manually entered miscellaneous tickets.wang-dsn17-pdf
Memory as a share of all failuresnot stated3.06percent of failuresSame ticketsPower 1.74%, RAID cards 1.23%.wang-dsn17-pdf
SSDs as a share of all failuresnot stated0.31percent of failuresSame ticketsFlash cards 0.67%.wang-dsn17-pdf
Failures in out-of-warranty hardware that are not handlednot stated25percent of failures (lower bound)Same ticketsPrinted as over 1/4 of failures; operators decommission the totally broken servers.wang-dsn17-pdf
Share of failed servers that produced over 99% of all failuresnot stated2percent of servers that ever failedSame ticketsPrinted as 2% of servers that ever failed contribute more than 99% of all failures.wang-dsn17-pdf
Fixed components that never repeat the same failurenot stated85percent (lower bound)Same ticketsPrinted as over 85%; about 4.5% of servers that ever failed had repeating failures.wang-dsn17-pdf
Days with over 500 hard drive failuresnot stated2.48percent of days35 of 1,411 daysBatch failures, counted per day.wang-dsn17-pdf
Servers of one product line reporting hard drive SMART alerts in a single nightnot stated32percent of that product line's serversNovember 2015 case study; 99% detected within about 6 hoursCause not identified; about 28% of the drives were replaced and over 70% were out of warranty.wang-dsn17-pdf
Mean time to respond to a failure ticketnot stated42.2daysTickets that led to a repair orderThe median is 6.1 days; this is operator response time, not a drive replacement time.wang-dsn17-pdf
Disks as a share of all hardware replacements2050percent of replacementsThree systems: HPC1 30%, COM2 50%, COM1 nearly 20%Replacement logs of other organizations, FAST 2007.schroeder-fast07-pdf
Node failures that are part of a burst of at least 2 nodesnot stated37percent of failuresTens of Google storage cells, 120 second windowPrinted as 37%; a different unit from servers with tickets.ford-osdi10-pdf
Older data blocks that fail a checksum on scrubbing0.00000010.000001fraction of older blocksGoogle file system, background scrubbingPrinted as 1 in 10^6 to 10^7 and concentrated on a small number of disks.ford-osdi10-pdf

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'over' or 'less than' it is the stated bound.
unit (string)
Unit of the value columns: tickets, years, percent, days, or fraction of blocks.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Wang et al. (DSN 2017), Schroeder and Gibson (FAST 2007), and Ford et al. (OSDI 2010), with credit. Not a Creative Commons license. The papers remain with their authors or publishers. This file is not a copy of the papers and not the data.

Sources

  1. What Can We Learn from Four Years of Data Center Hardware Failures?, Wang, Zhang, and Xu, IEEE/IFIP DSN 2017, pages 25 to 36 (author PDF, copy on the netman.aiops.org reading list) (accessed 2026-10-11)
  2. Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4 on other components) (accessed 2026-10-11)
  3. Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF) (accessed 2026-10-11)