Dataset · Google production fleet · December 2005–August 2006
Over 56% of failed drives had no count on four strong SMART signals
Over 56% of failed drives in the Google fleet studied by Pinheiro, Weber, and Barroso (FAST '07; more than 100,000 consumer PATA and SATA drives; SMART counts from December 2005 through August 2006) had no count on scan errors, reallocation count, offline reallocation, or probational count. Adding every other SMART parameter except temperature still left over 36% of failed drives at zero on all of those variables. That 56% figure is an upper bound on coverage for the four signals, not a false-positive-aware accuracy.
Download CSV12 rows · 19 columns · CSVMethod
The CSV transcribes sentences from Pinheiro, Weber, and Barroso, FAST ’07, proceedings pages 17–28. It does not digitize a chart and it does not refit a rate. The fleet is the Google production fleet in that paper: more than 100,000 consumer serial and parallel ATA drives, 5,400–7,200 rpm, 80–400 GB, at least nine models, commissioned in or after 2001. SMART counts were collected from December 2005 through August 2006. A failure is a drive replaced in a repair. That definition excludes upgrades. The parameters were not used in the repair diagnostics when the counts were collected.
Prevalence (%) is a share of drives. An AFR column is a factor versus drives with a zero count, not a percent and not a line rate. The 60-day column is a different factor: failure within 60 days after the first count, again versus zero-count drives. The paper reports a critical threshold when a count raises that 60-day probability by at least 10 times and the confidence is above 95%. For scan errors, reallocation count, and probational count, that threshold is one. Offline reallocation is reported after the first count, over 21 times, without the words “critical threshold” on that sentence. Empty cells are figures the source did not state.
The coverage rows keep the paper’s denominator. Over 56% and over 36% divide failed drives. The next sentence, an arbitrary rule of more than half the observed time above 40C, says about 36% of all drives. Those are separate rows. The heise row is a news restatement dated 16 February 2007, not a second measurement. Age-curve failure rates in the paper come from a longer repairs database and are not rows here.
Limits
Population and date: one Google production fleet, more than 100,000 consumer PATA and SATA drives (5,400–7,200 rpm, 80–400 GB, at least nine models, commissioned in or after 2001), SMART counts from December 2005 through August 2006. Not measured in this file: a false-positive rate for any predictor on this fleet, a per-model or per-manufacturer table, vibration, and the raw SMART or repairs logs. Figure 14 is not digitized. The age-curve rates from the separate repairs database are not rows.
- The over-56% miss is an upper bound on failed-drive coverage for four signals. Models that also keep false positives small are stated to do much worse. The paper does not print that false-positive rate.
- Manufacturer and model breakdowns are withheld as proprietary. Offline reallocation is not comparable across models, because some models report more offline reallocations than total reallocations. Seek error rate is the only SMART result the paper says changes with model mix, and only one manufacturer shows it widely.
- The two 36% sentences do not share a denominator. Figure 14’s axis is percent of the failed population. The Any Error bar crosses the 60% gridline, and the Seek Errors bar stays under the 40% gridline. This file does not turn those bars into percents.
- heise states 36 percent of failed drives, as a point, for no SMART error at all. That is not “over 36%” and not the four-signal cut. Spin retries contribute a count of zero, so they mark no failure in the window.
Comparison
| Signal | Population | Window | Drives with a non-zero count (%) | Bound on that share | AFR vs zero-count drives (factor) | Failure within 60 days vs zero-count drives (factor) | Bound on the 60-day factor | No-count share (%) | Bound on that share | No-count denominator | Signaled group that failed (%) | Bound on that share | Fault | What a count catches | What the signal misses | Citation |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Scan errors | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 2 | fewer-than | 10 | 39 | point | not stated | not stated | not stated | not stated | not stated | Surface defects found by a background scan of the media. | A non-zero count. The critical threshold is one, and the higher failure rate remains when the fleet is split by disk model. Non-zero counts are nearly uniform across models. Critical thresholds in the paper are counts that raise 60-day failure probability by at least 10 times versus zero-count drives, reported when confidence is above 95%. | A little over 70% of drives survive the first 8 months after the first scan error. The survival curve shows a 95% confidence interval. Drives with a zero count still fail. Together with the other three strong signals, these counts are still absent on over 56% of failed drives. The 10x figure is an AFR ratio, not the 39x 60-day ratio, and not a false-positive rate. | pinheiro-fast07 |
| Reallocation count | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 9 | about | 3–6 | 14 | over | not stated | not stated | not stated | not stated | not stated | A sector the drive treats as damaged, usually after recurring soft errors or a hard error, remapped to a spare. | A non-zero count, read as surface wear. The critical threshold is one. The AFR effect shows in every age group, and the trend is similar across models even where absolute counts differ. About 85% of drives survive past 8 months after the first reallocation. | The 3-6x AFR factor is weaker than the scan-error factor. The survival curve is cut at 8.5 months because confidence falls. The paper prints no per-model rate. A failed drive can still show a zero reallocation count. | pinheiro-fast07 |
| Offline reallocation | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 4 | about | not stated | 21 | over | not stated | not stated | not stated | not stated | not stated | A sector reallocated after background scrubbing. Defined as a subset of reallocation counts, leaving out sectors found during host I/O. | A non-zero count. After the first offline reallocation, the 60-day failure chance is over 21 times the chance for drives with none, a larger short-window effect than total reallocations. | Some models report more offline reallocations than total reallocations, so the count is not comparable across models. Non-zero counts concentrate on a subset of models. No prose AFR factor is stated. Figure 9 is not digitized. The paper cannot separate an older-age effect from the model mix. | pinheiro-fast07 |
| Probational count | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 2 | about | not stated | 16 | point | not stated | not stated | not stated | not stated | not stated | A suspect sector held on probation until it is reallocated or keeps working. | A non-zero count. The critical threshold is one. The paper calls it a softer error that might warn earlier. | A sector on probation may never be reallocated. The non-zero share is lower than both reallocation counts, which the paper reads as sectors leaving probation. Counts skew toward a subset of models. No prose AFR factor is stated. | pinheiro-fast07 |
| Four strong signals, union | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | not stated | not stated | not stated | not stated | not stated | 56 | over | failed drives | not stated | not stated | Failure with no count on scan errors, reallocation count, offline reallocation, or probational count. | Any non-zero count on those four. The paper singles them out as the SMART parameters with a large impact on failure probability. | Models that use only these signals can never predict more than half of the failed drives. The over-56% miss is an upper bound on failed-drive coverage in this fleet. It is not a false-positive-aware accuracy and not a per-model table. The paper says useful models that also keep false positives small are likely to do much worse. | pinheiro-fast07 |
| All SMART parameters except temperature | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | not stated | not stated | not stated | not stated | not stated | 36 | over | failed drives | not stated | not stated | Failure with every SMART count at zero, temperature left out. | Any remaining SMART variable, including seek error rate. Figure 14 is the percentage of failed drives with SMART errors. The plotted axis is percent of the failed population. Bar heights are not copied into this file. | Over 36% of failed drives still had zero counts on all of those variables. The same paragraph's figure of more than 72% of drives with seek errors is a fleet share, not this failed-drive share. The next sentence switches the denominator from failed drives to all drives. | pinheiro-fast07 |
| More than half the observed time above 40C | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | not stated | not stated | not stated | not stated | not stated | 36 | about | all drives | not stated | not stated | An arbitrary temperature rule, not a SMART error count. The rule is more than 50% of observed time above 40C. | Drives that meet the rule are added to the set the paper calls predictable failures. The paper says temperature has no crisp SMART error threshold. | After that addition the paper says about 36% of all drives still have no failure signal. The previous sentence says over 36% of failed drives had every count at zero. Those denominators are not the same, and the paper does not reconcile them. On Figure 14 the Any Error bar extends above the 60% gridline of the failed population, which fits a failed-drive zero-count share over 36%, not 36% of the fleet. This file does not record a bar height. | pinheiro-fast07 |
| CRC errors | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 2 | about | not stated | not stated | not stated | not stated | not stated | not stated | not stated | not stated | A CRC mismatch on the path between the physical media and the interface. | A non-zero count. The paper sees some correlation between higher CRC counts and failure, and calls that effect less pronounced than the strong signals. | CRC errors are less indicative of drive failure than of cables and connectors. No AFR factor and no 60-day factor are stated. A cable or connector fault can raise the count without a failed drive. | pinheiro-fast07 |
| Calibration retries | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 0.3 | under | not stated | not stated | not stated | not stated | not stated | not stated | 2 | about | Unspecified. The paper could not get one consistent definition from public documents or from some manufacturers. | A non-zero count. The parameter is present on a very small slice of the fleet. | Of the drives that have the count, only about 2% failed. The paper calls this a very weak and imprecise signal next to the other SMART parameters. A non-zero count does not name a fault in this study. | pinheiro-fast07 |
| Seek error rate | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 72 | more-than | not stated | not stated | not stated | not stated | not stated | not stated | not stated | not stated | The head failed to track a sector and the drive waited another revolution. | A rate, meant to be read against model-specific thresholds. Section 3.5.6 says this spread shrinks the set of drives with no error count at all. | The more-than-72% figure is a share of drives, not of failed drives. Seek errors are widespread for one manufacturer only, and that manufacturer's trend changes by vintage. For the other manufacturers the paper finds no correlation between failure rates and seek errors. On Figure 14 the Seek Errors bar stays below the 40% gridline of the failed population. This file does not record a bar height. Seek error rate is the only SMART result the paper says is changed by model mix. | pinheiro-fast07 |
| Spin retries | Google production fleet. More than 100,000 consumer serial and parallel ATA drives, 5400-7200 rpm, 80-400 GB, at least nine models, commissioned in or after 2001. A failure is a drive replaced in a repair. Upgrades are excluded. Manufacturer and model breakdowns are withheld as proprietary. | 2005-12/2006-08 | 0 | point | not stated | not stated | not stated | not stated | not stated | not stated | not stated | not stated | A retry while the drive attempts to spin up. | The parameter counts those spin-up retries. | The study registered not a single count in the whole population, so the signal marks no failed drive in this window. | pinheiro-fast07 |
| No SMART error before failure (news restatement) | News restatement of the same Google study, heise online, 16 February 2007. More than 100,000 IDE and SATA drives, 5400 and 7200 rpm, 80-400 GB, in service from 2001, watched for about nine months. Not a second measurement and not a per-model table. | about nine months | not stated | not stated | not stated | not stated | not stated | 36 | point | failed drives | not stated | not stated | A drive that failed during collection with no SMART error beforehand. The article's example is a burned chip on the drive electronics. | The article says scan errors indicate media-surface damage, and that a high reallocated-sector count usually led to drive death. It states no 60-day factor. | The article's 36 percent is a point for failed drives with no SMART error at all. It is not the paper's over 36%, and it is not the four-signal miss of over 56%. The article does not print the December 2005 through August 2006 bounds, a false-positive rate, or a model table. | heise-2007-02-16 |
The four-signal union is the coverage result: over 56% of failed drives, not a share of the fleet. Scan errors are fewer than 2% of drives, with an AFR factor of 10 and a 60-day factor of 39. Reallocation count is about 9% of drives, AFR factor 3–6, 60-day factor over 14. Those two factors are not interchangeable.
Columns
- row_id
- Stable id for the signal or coverage cut.
- signal
- SMART parameter, coverage cut, or the heise restatement of a coverage cut.
- population
- Fleet, media, and failure definition the row's figures belong to. Google rows are one production fleet. The heise row is a news restatement of that study, not a second fleet.
- window
- Observation window. Google SMART rows are December 2005 through August 2006 (2005-12/2006-08). The heise row states only about nine months.
- prevalence_percent
- Share of drives with a non-zero count, in percent. The number is the figure printed next to the bound in prevalence_bound, not a refit. Empty when the row is not a prevalence.
- prevalence_bound
- How prevalence_percent was stated: fewer-than, under, about, more-than, or point. Empty when prevalence is empty.
- afr_factor_low
- Low end of the stated AFR multiple versus drives with a zero count. A factor, not a percent and not a line rate. Empty when the paper states no AFR factor. Equal to afr_factor_high when the paper states one factor.
- afr_factor_high
- High end of that AFR multiple. Empty when afr_factor_low is empty.
- fail_60d_factor
- Multiple for failure within 60 days after the first count, versus drives with a zero count. Not an AFR. The paper reports a critical threshold when a count raises that probability by at least 10 times at confidence above 95%. Empty when no 60-day factor is stated.
- fail_60d_bound
- How fail_60d_factor was stated: point or over. Empty when the factor is empty.
- no_count_share_percent
- Percent of the denominator in no_count_denominator that had no failure signal for this cut. Empty when the row is not a coverage sentence.
- no_count_share_bound
- How no_count_share_percent was stated: over, about, or point. Empty when the share is empty.
- no_count_denominator
- Who the no-count share divides: failed-drives or all-drives. The paper uses both for a 36% figure. Empty when the share is empty.
- signaled_group_failed_percent
- Percent of drives that already had the count and then failed. Used for calibration retries only. Not an AFR and not a share of all failed drives. Empty otherwise.
- signaled_group_failed_bound
- How signaled_group_failed_percent was stated: about or point. Empty when that percent is empty.
- fault
- Named condition the signal is about, from a sentence in the cited source.
- catches
- What a non-zero count indicates in that source.
- misses
- The failure or ambiguity that count does not prove, from a sentence in the same source.
- citation_id
- Id into the citations list. Open that URL for the sentence.
License
Small derived table of figures stated in Pinheiro, Weber, and Barroso, Failure Trends in a Large Disk Drive Population, FAST '07, proceedings pages 17-28, and of one restatement in heise online (16 February 2007). Not a Creative Commons license. USENIX's open-access notice says papers and proceedings are freely available to everyone once the event begins. That notice is not a CC grant, and the FAST '07 HTML does not grant one. This CSV is not Google's SMART repository or repairs database. Canonical files are the USENIX HTML paper and the proceedings PDF.
USENIX open-access notice · Canonical FAST ’07 HTML · Proceedings PDF (not mirrored here) · heise online, 16 February 2007