Dataset · silent corruption · FAST 2008
0.86% of nearline disks and 0.065% of enterprise disks developed a checksum mismatch in 41 months
Bairavasundaram et al., FAST 2008, report that 3,088 of about 358,000 nearline disks (0.86%) and 767 of 1.17 million enterprise-class disks (0.065%) developed at least one silent checksum mismatch in 41 months of support logs covering 1.53 million disks, about 400,000 mismatched 4 KB blocks in all. The median affected disk had 3 mismatches, the mean 104, the maximum 33,000. Only detected corruption is counted. A CERN IT report from 2007 gives an independent but different measure: 22 of 33,700 files failed a checksum comparison.
Download CSV18 rows · 7 columns · CSVMethod
The CSV transcribes figures printed in Bairavasundaram et al. (FAST 2008) and two counts from a 2007 CERN IT report. It does not read chart values. A checksum mismatch is a 4 KB block that fails its stored checksum when the RAID layer reads it. Percent of disks is the share of disks with at least one event over the stated window. Numerals the HTML conversion of the paper drops were checked against the PDF.
Limits
- Only corruption the storage system detected is counted. Corruption that is never read or scrubbed is invisible.
- Logging is optional for customers and removed disks are counted only until removal.
- A disk's result includes its adapter and shelf path. The paper names faulty adapters or shelf controllers for three models.
- The paper prints a mean of 104 for the whole sample and 78 in the first-17-months analysis.
- The CERN report is a 2007 internal draft from one site with its own probes. Its counts are not per-disk shares and do not test the NetApp percentages.
- Several results appear only as charts in the paper and are not in this file.
Table
| Measure | Value low | Value high | Unit | Scope | Note | Citation |
|---|---|---|---|---|---|---|
| disks in the sample | not stated | 1530000 | disks | 41 months from January 2004; 14 families; 31 models | About 358000 nearline and 1.17 million enterprise-class disks. | bairavasundaram-usenix-pdf |
| checksum mismatches observed | not stated | 400000 | 4 KB blocks | Whole sample for 41 months | Printed as about 400000. A checksum mismatch is a block that fails its stored checksum on read. | bairavasundaram-usenix-html |
| nearline disks with at least one mismatch | not stated | 0.86 | percent of disks | About 358000 nearline disks for 41 months | 3088 disks. | bairavasundaram-usenix-pdf |
| enterprise disks with at least one mismatch | not stated | 0.065 | percent of disks | 1.17 million enterprise-class disks for 41 months | 767 disks. | bairavasundaram-usenix-pdf |
| nearline disks with a mismatch in the first 17 months | not stated | 0.66 | percent of disks | Weighted nearline average | Detected corruption only. | bairavasundaram-usenix-pdf |
| enterprise disks with a mismatch in the first 17 months | not stated | 0.06 | percent of disks | Weighted enterprise average | Detected corruption only. | bairavasundaram-usenix-pdf |
| median mismatches per corrupt disk | not stated | 3 | blocks per disk | 3855 corrupt disks over 41 months | Mode is 1. | bairavasundaram-usenix-pdf |
| mean mismatches per corrupt disk | not stated | 104 | blocks per disk | 3855 corrupt disks over 41 months | The first-17-months analysis in the paper prints a mean of 78. | bairavasundaram-usenix-pdf |
| maximum mismatches on one disk | not stated | 33000 | blocks per disk | Whole sample | Single drive. | bairavasundaram-usenix-pdf |
| share of mismatches from the top 1 percent of corrupt disks | not stated | 50 | percent of mismatches | Corrupt disks only | Printed as more than half so 50 is a lower bound. | bairavasundaram-usenix-html |
| chance of further mismatches after one on a nearline disk | not stated | 0.6 | probability (unitless) | First 17 months in the field | Compare 0.0066 for a first mismatch. | bairavasundaram-usenix-html |
| mismatches found by scrubs on nearline disks | not stated | 49 | percent of mismatches | Weighted nearline average | Printed as about 49. | bairavasundaram-usenix-html |
| mismatches found by scrubs on enterprise disks | not stated | 73 | percent of mismatches | Weighted enterprise average | Printed as about 73. | bairavasundaram-usenix-html |
| mismatches found during RAID reconstruction on nearline disks | not stated | 8 | percent of mismatches | Weighted nearline average | Printed as about 8. | bairavasundaram-usenix-html |
| nearline disks affected per year by latent sector errors | not stated | 9.5 | percent of disks per year | Table 2 average | Compare 0.466 for checksum mismatches. | bairavasundaram-usenix-html |
| nearline disks affected per year by checksum mismatches | not stated | 0.466 | percent of disks per year | Table 2 average | Detected corruption only. | bairavasundaram-usenix-html |
| CERN files failing a checksum comparison | not stated | 22 | files | 33700 files (about 8.7 TB) on a CERN disk pool in early 2007 | About 1 bad file in 1500. A different measure from the per-disk shares. | cern-data-integrity-2007 |
| CERN RAID verification block problems | not stated | 300 | blocks in 4 weeks | 492 systems (about 1.5 PB) | The report says the vendor bit error rate would predict about 850. | cern-data-integrity-2007 |
Columns
- measure
- The quantity as the source names it.
- value_low
- Lower end of a printed range. Empty when the source prints a single value.
- value_high
- Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
- unit
- Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
- scope
- Population and window the value applies to.
- note
- What the value is not, or the source wording behind it.
- citation_id
- Id of the opened source in the citations list.
License
Small derived table of figures printed in Bairavasundaram et al. (FAST 2008) and the CERN IT report Data integrity (Panzer-Steindel, draft 1.3, 2007), with credit. Not a Creative Commons license. The USENIX proceedings state that rights to individual papers remain with the author or the author's employer. This file is not a copy of either document and not the underlying support logs or probe data.