Dataset · Hardware failures · DSN 2017
Hard drives were 82% of 290,000 datacenter hardware failure tickets, and 2% of failed servers produced over 99% of the failures
Wang, Zhang, and Xu, DSN 2017, analysed over 290,000 hardware failure tickets from four years at one large Internet company. Hard drives were 81.84% of failures and 2% of failed servers produced over 99% of failures. Cross-check rows come from Schroeder and Gibson (FAST 2007) and Ford et al. (OSDI 2010).
Download CSV14 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Wang, Zhang, and Xu (DSN 2017) and cross-check rows from Schroeder and Gibson (FAST 2007) and Ford et al. (OSDI 2010). It does not read chart values.
Limits
- One company's fleet and ticketing rules; the company is not named in the paper.
- A failure is a ticket and includes SMART and error warnings, so shares are not failure rates per drive.
- The 42.2 day figure is operator response time, not repair time.
- The cause of the November 2015 batch is not known.
- The cross-check rows are other systems with other definitions.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| Hardware failure tickets analysed | not stated | 290000 | tickets | One large Internet company, dozens of datacenters, hundreds of thousands of servers | Printed as over 290,000 reports. | wang-dsn17-pdf |
| Length of the record | not stated | 4 | years | Same company | Printed as the past four years; the batch analysis covers 1,411 days. | wang-dsn17-pdf |
| Hard drives as a share of all failures | not stated | 81.84 | percent of failures | Failure tickets excluding false alarms | Text says about 82%; 10.20% are manually entered miscellaneous tickets. | wang-dsn17-pdf |
| Memory as a share of all failures | not stated | 3.06 | percent of failures | Same tickets | Power 1.74%, RAID cards 1.23%. | wang-dsn17-pdf |
| SSDs as a share of all failures | not stated | 0.31 | percent of failures | Same tickets | Flash cards 0.67%. | wang-dsn17-pdf |
| Failures in out-of-warranty hardware that are not handled | not stated | 25 | percent of failures (lower bound) | Same tickets | Printed as over 1/4 of failures; operators decommission the totally broken servers. | wang-dsn17-pdf |
| Share of failed servers that produced over 99% of all failures | not stated | 2 | percent of servers that ever failed | Same tickets | Printed as 2% of servers that ever failed contribute more than 99% of all failures. | wang-dsn17-pdf |
| Fixed components that never repeat the same failure | not stated | 85 | percent (lower bound) | Same tickets | Printed as over 85%; about 4.5% of servers that ever failed had repeating failures. | wang-dsn17-pdf |
| Days with over 500 hard drive failures | not stated | 2.48 | percent of days | 35 of 1,411 days | Batch failures, counted per day. | wang-dsn17-pdf |
| Servers of one product line reporting hard drive SMART alerts in a single night | not stated | 32 | percent of that product line's servers | November 2015 case study; 99% detected within about 6 hours | Cause not identified; about 28% of the drives were replaced and over 70% were out of warranty. | wang-dsn17-pdf |
| Mean time to respond to a failure ticket | not stated | 42.2 | days | Tickets that led to a repair order | The median is 6.1 days; this is operator response time, not a drive replacement time. | wang-dsn17-pdf |
| Disks as a share of all hardware replacements | 20 | 50 | percent of replacements | Three systems: HPC1 30%, COM2 50%, COM1 nearly 20% | Replacement logs of other organizations, FAST 2007. | schroeder-fast07-pdf |
| Node failures that are part of a burst of at least 2 nodes | not stated | 37 | percent of failures | Tens of Google storage cells, 120 second window | Printed as 37%; a different unit from servers with tickets. | ford-osdi10-pdf |
| Older data blocks that fail a checksum on scrubbing | 0.0000001 | 0.000001 | fraction of older blocks | Google file system, background scrubbing | Printed as 1 in 10^6 to 10^7 and concentrated on a small number of disks. | ford-osdi10-pdf |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'over' or 'less than' it is the stated bound.
- unit
- Unit of the value columns: tickets, years, percent, days, or fraction of blocks.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Wang et al. (DSN 2017), Schroeder and Gibson (FAST 2007), and Ford et al. (OSDI 2010), with credit. Not a Creative Commons license. The papers remain with their authors or publishers. This file is not a copy of the papers and not the data.
Sources
- What Can We Learn from Four Years of Data Center Hardware Failures?, Wang, Zhang, and Xu, IEEE/IFIP DSN 2017, pages 25 to 36 (author PDF, copy on the netman.aiops.org reading list)
- Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4 on other components)
- Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF)