Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · silent corruption · FAST 2008

0.86% of nearline disks and 0.065% of enterprise disks developed a checksum mismatch in 41 months

Bairavasundaram et al., FAST 2008, report that 3,088 of about 358,000 nearline disks (0.86%) and 767 of 1.17 million enterprise-class disks (0.065%) developed at least one silent checksum mismatch in 41 months of support logs covering 1.53 million disks, about 400,000 mismatched 4 KB blocks in all. The median affected disk had 3 mismatches, the mean 104, the maximum 33,000. Only detected corruption is counted. A CERN IT report from 2007 gives an independent but different measure: 22 of 33,700 files failed a checksum comparison.

Download CSV18 rows · 7 columns · CSV

File /data/datasets/silent-corruption-checksum-mismatches.csv · JSON metadata · All datasets · Human-readable note on hesela.com

Method

The CSV transcribes figures printed in Bairavasundaram et al. (FAST 2008) and two counts from a 2007 CERN IT report. It does not read chart values. A checksum mismatch is a 4 KB block that fails its stored checksum when the RAID layer reads it. Percent of disks is the share of disks with at least one event over the stated window. Numerals the HTML conversion of the paper drops were checked against the PDF.

Limits

Table

18 rows, the same rows as the CSV. Value low and value high are the printed range ends, or a single printed value in the high column. Units are in the unit column. Sources are the pages opened on 2026-10-10.
MeasureValue lowValue highUnitScopeNoteCitation
disks in the samplenot stated1530000disks41 months from January 2004; 14 families; 31 modelsAbout 358000 nearline and 1.17 million enterprise-class disks.bairavasundaram-usenix-pdf
checksum mismatches observednot stated4000004 KB blocksWhole sample for 41 monthsPrinted as about 400000. A checksum mismatch is a block that fails its stored checksum on read.bairavasundaram-usenix-html
nearline disks with at least one mismatchnot stated0.86percent of disksAbout 358000 nearline disks for 41 months3088 disks.bairavasundaram-usenix-pdf
enterprise disks with at least one mismatchnot stated0.065percent of disks1.17 million enterprise-class disks for 41 months767 disks.bairavasundaram-usenix-pdf
nearline disks with a mismatch in the first 17 monthsnot stated0.66percent of disksWeighted nearline averageDetected corruption only.bairavasundaram-usenix-pdf
enterprise disks with a mismatch in the first 17 monthsnot stated0.06percent of disksWeighted enterprise averageDetected corruption only.bairavasundaram-usenix-pdf
median mismatches per corrupt disknot stated3blocks per disk3855 corrupt disks over 41 monthsMode is 1.bairavasundaram-usenix-pdf
mean mismatches per corrupt disknot stated104blocks per disk3855 corrupt disks over 41 monthsThe first-17-months analysis in the paper prints a mean of 78.bairavasundaram-usenix-pdf
maximum mismatches on one disknot stated33000blocks per diskWhole sampleSingle drive.bairavasundaram-usenix-pdf
share of mismatches from the top 1 percent of corrupt disksnot stated50percent of mismatchesCorrupt disks onlyPrinted as more than half so 50 is a lower bound.bairavasundaram-usenix-html
chance of further mismatches after one on a nearline disknot stated0.6probability (unitless)First 17 months in the fieldCompare 0.0066 for a first mismatch.bairavasundaram-usenix-html
mismatches found by scrubs on nearline disksnot stated49percent of mismatchesWeighted nearline averagePrinted as about 49.bairavasundaram-usenix-html
mismatches found by scrubs on enterprise disksnot stated73percent of mismatchesWeighted enterprise averagePrinted as about 73.bairavasundaram-usenix-html
mismatches found during RAID reconstruction on nearline disksnot stated8percent of mismatchesWeighted nearline averagePrinted as about 8.bairavasundaram-usenix-html
nearline disks affected per year by latent sector errorsnot stated9.5percent of disks per yearTable 2 averageCompare 0.466 for checksum mismatches.bairavasundaram-usenix-html
nearline disks affected per year by checksum mismatchesnot stated0.466percent of disks per yearTable 2 averageDetected corruption only.bairavasundaram-usenix-html
CERN files failing a checksum comparisonnot stated22files33700 files (about 8.7 TB) on a CERN disk pool in early 2007About 1 bad file in 1500. A different measure from the per-disk shares.cern-data-integrity-2007
CERN RAID verification block problemsnot stated300blocks in 4 weeks492 systems (about 1.5 PB)The report says the vendor bit error rate would predict about 850.cern-data-integrity-2007

Columns

measure (string)
The quantity as the source names it.
value_low (number)
Lower end of a printed range. Empty when the source prints a single value.
value_high (number)
Single printed value, or the upper end of a range. For 'more than half' it is the stated bound.
unit (string)
Unit of the value columns: disks, blocks, files, percent of disks, percent of mismatches, or probability.
scope (string)
Population and window the value applies to.
note (string)
What the value is not, or the source wording behind it.
citation_id (string)
Id of the opened source in the citations list.

License

Small derived table of figures printed in Bairavasundaram et al. (FAST 2008) and the CERN IT report Data integrity (Panzer-Steindel, draft 1.3, 2007), with credit. Not a Creative Commons license. The USENIX proceedings state that rights to individual papers remain with the author or the author's employer. This file is not a copy of either document and not the underlying support logs or probe data.

Sources

  1. An Analysis of Data Corruption in the Storage Stack, Bairavasundaram et al., FAST 2008 (USENIX HTML proceedings) (accessed 2026-10-10)
  2. An Analysis of Data Corruption in the Storage Stack (USENIX PDF, used for numerals the HTML conversion drops) (accessed 2026-10-10)
  3. Data integrity, Bernd Panzer-Steindel, CERN IT, draft 1.3, 8 April 2007 (accessed 2026-10-10)