Dataset · vendor telemetry · two NetApp field studies
Low-end NetApp interconnect events exceed disk events, 4,338 to 3,230
In Jiang et al., FAST 2008, Table 1, low-end NetApp systems recorded 4,338 physical-interconnect failure events and 3,230 disk failure events from January 2004 through August 2007 (22,031 systems and 264,983 FC disks ever installed). For Figure 4(b), which excludes subsystems that used Disk H, that class has a stated disk AFR of 0.9% and a storage-subsystem AFR of 4.6%. In Maneas et al., FAST 2020, Table 2, SCSI error is 32.78% of SSD replacements (ARR 0.055%) on almost 1.4 million NetApp SSDs; that signal is a hardware error the SSD reports (the example is drive DRAM ECC), while aborted commands are 13.56% (ARR 0.023%) and include a host or connection case where write data never reaches the device.
Download CSV33 rows · 18 columns · CSVMethod
The CSV transcribes stated figures. It does not digitize a chart and it does not refit a rate. Jiang et al. report NetApp AutoSupport logs from about 39,000 commercially deployed storage systems over 44 months, January 2004 through August 2007, with about 1,800,000 disks in about 155,000 shelf enclosures. Table 1 prints the class counts and the four failure-event counts. Those event rows are the jiang-table1 cut. The disk column is disks ever installed during the 44 months, including replacements, not a one-day census.
Figure 4(b) excludes subsystems that used the problematic Disk H family. The AFR cells are the figures written in that section, including Finding (2): near-line disk AFR 1.9%, near-line subsystem AFR 3.4%, low-end disk AFR 0.9%, low-end subsystem AFR 4.6%. Nearby sentences also say “about” for the 1.9%, 3.4%, and 4.6% figures. The low-end example says disks are about 20% of that subsystem AFR. Those four rows are jiang-figure4b. They are not Table 1 events, and this file does not divide 0.9 by 4.6.
The share rows are the cross-class prose ranges: disk failures 20–55% of storage-subsystem failures, physical interconnects including shelf enclosures 27–68%, protocol failures 5–10%, and performance failures 4–8%. They are not percentages of the Table 1 event counts. Annual rate is AFR only on the Figure 4(b) rows. ARR is a different definition, failures divided by device-years, and appears only on the Maneas rows.
Maneas et al. use ten NetApp Active IQ snapshots (previously AutoSupport), January and June 2017, January, May, August, and December 2018, and February, March, April, and May 2019: almost 1.4 million enterprise SSDs, three manufacturers, 18 models, over 30 months. Table 2 prints nine replacement reasons. The reason was missing for 40% of replacements, and the percent of replacements is normalized for that gap. SCSI error is a hardware error the SSD reports; the paper’s example is drive DRAM ECC. Aborted commands are the host-or-connection case. There is no line-rate or throughput column in this file.
Limits
- Population and date, HDD side: NetApp customer AutoSupport, four system classes, January 2004 through August 2007. Not an industry sample. The counts are the paper’s figures from that vendor telemetry, not a recount of the logs.
- Population and date, SSD side: a sample of NetApp Active IQ, almost 1.4 million SSDs, ten snapshots spanning January 2017 through May 2019 (the paper says 30 months). Not a daily failure log, and not the HDD population.
- Not measured here: mid-range and high-end point disk AFR, and mid-range and high-end subsystem AFR. The prose says disk AFR for low-end, mid-range, and high-end together is under 0.9%, and separately states the low-end point as 0.9%. This file does not pick a point for the other two classes. Figure 4(a), which still includes Disk H, is not transcribed.
- Not measured here: failures that never reach the RAID layer. The paper’s example is an interconnect failure recovered by SCSI retries or tolerated by multipathing. Higher-layer redundancy above the storage subsystem is outside the 2008 study.
- Not measured here: raw SSD replacement counts, a per-reason confidence interval, and the missing reason on 40% of SSD replacements. The bundles do not contain copies of customers’ data. Lost writes stay ambiguous between drive firmware and another layer in the stack.
- Table 1 event shares are not the 20–55% or 27–68% ranges. Those ranges are the paper’s failure-composition prose. Do not treat the summed ARR cells as a published fleet ARR. Neither paper links a public dump.
Counted here
From the Table 1 rows, not from the paper’s rounded “about” lines: 39,115 systems, 156,990 shelf enclosures, 1,819,423 disks ever installed. The paper says about 39,000 systems, about 155,000 shelf enclosures, and about 1,800,000 disks. This file does not force those sentences to equal the sums.
- near-line: 17,892 failure events across the four Table 1 types.
- low-end: 9,824 failure events across the four Table 1 types.
- mid-range: 21,296 failure events across the four Table 1 types.
- high-end: 17,364 failure events across the four Table 1 types.
Low-end physical-interconnect events (4,338) exceed low-end disk events (3,230). Near-line disk events are 10,105 against 4,888 physical-interconnect events, so the low-end ordering is not the ordering in every class.
The nine Table 2 share cells sum to 100. Predictive failures, threshold exceeded, and recommended failures sum to 34.44% of replacements. The paper calls that preventative group about one third and does not print 34.44. The nine ARR cells sum to 0.166. Table 2 does not print that sum, so this file does not call 0.166% the fleet annual replacement rate.
Comparison
| Cut | Study | Fleet | Window | Class | Population | Fault | Systems (count) | Shelf enclosures (count) | Disks ever installed (count) | Failure events (count) | Share, low (%) | Share, high (%) | Annual rate (%) | Annual-rate kind | Signal | What the signal misses | Citation |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | near-line | Table 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | disk failure | 4,927 | 33,681 | 520,776 | 10,105 | not stated | not stated | not stated | not stated | RAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors. | Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | near-line | Table 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | physical interconnect failure | 4,927 | 33,681 | 520,776 | 4,888 | not stated | not stated | not stated | not stated | Disks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error. | An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | near-line | Table 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | protocol failure | 4,927 | 33,681 | 520,776 | 1,819 | not stated | not stated | not stated | not stated | Disks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers. | The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | near-line | Table 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | performance failure | 4,927 | 33,681 | 520,776 | 1,080 | not stated | not stated | not stated | not stated | The storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping. | The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | low-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | disk failure | 22,031 | 37,260 | 264,983 | 3,230 | not stated | not stated | not stated | not stated | RAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors. | Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | low-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | physical interconnect failure | 22,031 | 37,260 | 264,983 | 4,338 | not stated | not stated | not stated | not stated | Disks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error. | An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | low-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | protocol failure | 22,031 | 37,260 | 264,983 | 1,021 | not stated | not stated | not stated | not stated | Disks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers. | The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | low-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | performance failure | 22,031 | 37,260 | 264,983 | 1,235 | not stated | not stated | not stated | not stated | The storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping. | The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | mid-range | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | disk failure | 7,154 | 52,621 | 578,980 | 8,989 | not stated | not stated | not stated | not stated | RAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors. | Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | mid-range | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | physical interconnect failure | 7,154 | 52,621 | 578,980 | 7,949 | not stated | not stated | not stated | not stated | Disks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error. | An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | mid-range | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | protocol failure | 7,154 | 52,621 | 578,980 | 2,298 | not stated | not stated | not stated | not stated | Disks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers. | The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | mid-range | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | performance failure | 7,154 | 52,621 | 578,980 | 2,060 | not stated | not stated | not stated | not stated | The storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping. | The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | high-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | disk failure | 5,003 | 33,428 | 454,684 | 8,240 | not stated | not stated | not stated | not stated | RAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors. | Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | high-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | physical interconnect failure | 5,003 | 33,428 | 454,684 | 7,395 | not stated | not stated | not stated | not stated | Disks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error. | An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | high-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | protocol failure | 5,003 | 33,428 | 454,684 | 1,576 | not stated | not stated | not stated | not stated | Disks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers. | The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count. | jiang-fast08 |
| jiang-table1 | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | high-end | Table 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census. | performance failure | 5,003 | 33,428 | 454,684 | 153 | not stated | not stated | not stated | not stated | The storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping. | The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count. | jiang-fast08 |
| jiang-figure4b | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | near-line | Figure 4(b) prose for near-line systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count. | disk failure | not stated | not stated | not stated | not stated | not stated | not stated | 1.9 | AFR | Finding (2) states a 1.9% disk AFR. The next paragraph, still on Figure 4(b), says near-line systems mostly use SATA disks and experience about 1.9% AFR for disks. | This row excludes subsystems that used Disk H. It is not the Table 1 event count. The prose does not give a near-line disk share of subsystem AFR. | jiang-fast08 |
| jiang-figure4b | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | near-line | Figure 4(b) prose for near-line systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count. | storage subsystem | not stated | not stated | not stated | not stated | not stated | not stated | 3.4 | AFR | Finding (2) states a 3.4% storage-subsystem AFR. A later sentence on the same figure says the near-line storage-subsystem AFR is about 3.4%, lower than the low-end figure. | Disk H is excluded. The paper does not print this AFR as a sum of Table 1 events. | jiang-fast08 |
| jiang-figure4b | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | low-end | Figure 4(b) prose for low-end systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count. | disk failure | not stated | not stated | not stated | not stated | 20 | 20 | 0.9 | AFR | The Figure 4(b) example says the disk AFR is only 0.9%, about 20% of overall subsystem AFR. Finding (2) states 0.9%. A later sentence says disk AFR for low-end, mid-range, and high-end together is under 0.9%. | The 20% cell is the paper's about-20% of subsystem AFR on this Disk H-excluded cut, not a share of Table 1 events. This file does not give mid-range or high-end a point disk AFR. | jiang-fast08 |
| jiang-figure4b | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | low-end | Figure 4(b) prose for low-end systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count. | storage subsystem | not stated | not stated | not stated | not stated | not stated | not stated | 4.6 | AFR | The Figure 4(b) example says the storage-subsystem AFR is about 4.6%. Finding (2) states 4.6%, higher than the near-line subsystem AFR. | Disk H is excluded. Not a Table 1 event sum. The 0.9% disk AFR is a different row. | jiang-fast08 |
| jiang-share | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | all four classes | Prose across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4. | disk failure | not stated | not stated | not stated | not stated | 20 | 55 | not stated | not stated | Disk failures contribute to 20-55% of storage subsystem failures. | Cross-class range in prose. Not the low-end about-20% point, and not a percentage of Table 1 events. | jiang-fast08 |
| jiang-share | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | all four classes | Prose across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4. | physical interconnect failure | not stated | not stated | not stated | not stated | 27 | 68 | not stated | not stated | Physical interconnects, including shelf enclosures, account for 27-68% of storage subsystem failures. The Figure 4(b) discussion says that fraction ranges from 27% to 68%. | Cross-class prose range. Not a Table 1 event share. An interconnect fault recovered by SCSI retries or tolerated by multipathing is not in the subsystem count. | jiang-fast08 |
| jiang-share | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | all four classes | Prose across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4. | protocol failure | not stated | not stated | not stated | not stated | 5 | 10 | not stated | not stated | Protocol failures contribute to 5-10% of storage subsystem failures. | Cross-class prose range. The tag means disks are visible but I/O is not correctly answered. | jiang-fast08 |
| jiang-share | jiang-fast08 | NetApp AutoSupport | 2004-01/2007-08 | all four classes | Prose across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4. | performance failure | not stated | not stated | not stated | not stated | 4 | 8 | not stated | not stated | Performance failures contribute to 4-8% of storage subsystem failures. | Cross-class prose range. Counted when the disk does not serve I/O in time and the other three types were not detected. | jiang-fast08 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | SCSI error | not stated | not stated | not stated | not stated | 32.78 | 32.78 | 0.055 | ARR | The SCSI layer reports a hardware error from the SSD that is severe enough to replace the drive and reconstruct the data. The example is ECC errors from the drive's DRAM. | Normalized because the reason was missing for 40% of replacements. This is not the aborted-command row. The paper's host or connection case, where write data never reaches the device, is that other row. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | unresponsive drive | not stated | not stated | not stated | not stated | 0.60 | 0.60 | 0.001 | ARR | The drive has completely failed and become unresponsive. | Normalized share, not a raw ticket count. The prose says this reason is reported for only 0.60% of replacements. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | lost writes | not stated | not stated | not stated | not stated | 13.54 | 13.54 | 0.023 | ARR | A 4K WAFL block read from the SSD is inconsistent with its signature, which includes attributes and a version number. The replace heuristic is multiple such errors on one SSD and none on any other SSD. | The paper says the root cause could be a firmware bug in the drive or another layer in the stack, and that it is less clear whether the drive or other layers are to blame. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | aborted commands | not stated | not stated | not stated | not stated | 13.56 | 13.56 | 0.023 | ARR | An aborted command reported either by the SSD or by the storage layer. | The stated case is a host or connection issue: the host sends write commands, but the data never reach the device. The tag does not by itself prove an SSD hardware fault. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | disk-ownership I/O errors | not stated | not stated | not stated | not stated | 3.27 | 3.27 | 0.005 | ARR | An error while communicating with the subsystem that tracks which node owns a disk. The SSD is then marked failed. | Normalized share. The paper ties the tag to that ownership subsystem. It does not give a NAND or DRAM example for this row. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | command timeouts | not stated | not stated | not stated | not stated | 1.81 | 1.81 | 0.003 | ARR | The SSD's own timers, and the storage layer's timers, expire. The operation does not finish in the allotted time even after retries. | Normalized share. The definition does not identify the timeout as media wear-out. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | predictive failures | not stated | not stated | not stated | not stated | 12.78 | 12.78 | 0.021 | ARR | The SSD reports a pattern of recovered errors using the manufacturer's own thresholds and criteria. | Preventative: the prose says category D drives were still operational before replacement. A further split of predictive replacements is not in Table 2. The paper says the most common trigger behind a preventative replacement is exceeding the threshold of consecutive timeouts. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | threshold exceeded | not stated | not stated | not stated | not stated | 12.73 | 12.73 | 0.020 | ARR | The Storage Health Monitor replaces the SSD when a threshold is crossed. The example is the number of media errors. | Preventative. The drive was still operational. Recommended failures are described as less strict and less urgent than this row. | maneas-fast20 |
| maneas-table2 | maneas-fast20 | NetApp Active IQ | 2017-01/2019-05 | ssd sample | Table 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count. | recommended failures | not stated | not stated | not stated | not stated | 8.93 | 8.93 | 0.015 | ARR | The system reports that the drive should be replaced in the near future. | Less strict and less urgent than threshold exceeded. The drive was still operational before replacement. | maneas-fast20 |
Window, fleet, and population are columns in the CSV. Jiang rows use NetApp AutoSupport, 2004-01/2007-08. Maneas rows use NetApp Active IQ, 2017-01/2019-05. The population cell states which cut the numbers belong to.
Columns
- row_kind
- Which stated cut the row transcribes: jiang-table1, jiang-figure4b, jiang-share, or maneas-table2.
- study
- Citation id of the paper that states the row.
- fleet
- Named fleet: NetApp AutoSupport or NetApp Active IQ.
- window
- Study window as year-month bounds. Jiang is January 2004 through August 2007 (44 months). Maneas is the 30 months covered by 10 snapshots from January 2017 through May 2019.
- system_class
- Jiang storage-system class (near-line, low-end, mid-range, high-end, or all four classes) or ssd sample.
- population
- Who was counted, which cut (Table 1, Figure 4(b) excluding Disk H, cross-class prose, or Table 2), and what the count is not.
- fault
- Named failure or replacement reason, using the paper's categories.
- systems_count
- Systems (count) from Jiang Table 1. Empty when the row is not a Table 1 class row. Not a Figure 4(b) population.
- shelves_count
- Shelf enclosures (count) from Jiang Table 1. Empty when the row is not a Table 1 class row.
- disks_ever_installed_count
- Disks ever installed (count) during the 44 months, including replacements, from Jiang Table 1. Empty when the row is not a Table 1 class row. Not the Maneas SSD population.
- failure_events
- Failure events (count) for that fault in Jiang Table 1. Empty when the source did not state an event count for the row.
- share_low_percent
- Share, low (%). Jiang cross-class prose uses a range. The low-end Figure 4(b) disk row uses 20 for the paper's about-20% of subsystem AFR. Maneas Table 2 percent of replacements is a point, so low equals high. Empty when not stated.
- share_high_percent
- Share, high (%). Equal to share_low_percent when the source stated a point rather than a range. Empty when not stated.
- annual_rate_percent
- Annual rate (%). AFR for Jiang Figure 4(b) prose, or ARR for Maneas Table 2. Empty when the source did not state an annual rate for the row. Not a line rate and not a throughput.
- annual_rate_kind
- AFR or ARR when annual_rate_percent is set. Empty otherwise. AFR and ARR are different definitions and are not converted into each other.
- signal
- The log or telemetry signal the paper says catches this fault.
- misses
- What that signal does not prove, from a sentence in the same paper.
- citation_id
- Id into the citations list. Open that URL for the sentence.
License
Small derived table of figures stated in Jiang et al., FAST 2008, and Maneas et al., FAST 2020. Not a Creative Commons license. The opened FAST '14 Consent Form for Refereed Papers says USENIX lets authors retain copyright and asks only for a right to publish. That grant is non-exclusive, worldwide, perpetual, and irrevocable, and also non-sublicensable and non-transferable. The form contains no CC grant. This CSV is not the NetApp AutoSupport database and not the Active IQ bundles. Neither paper links a public dump or a data DOI for those logs. Canonical files are the USENIX PDFs.
FAST ’14 refereed-paper consent form · Canonical FAST 2008 PDF (AutoSupport logs are not in the file) · Canonical FAST 2020 PDF (Active IQ bundles are not in the file)