Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · vendor telemetry · two NetApp field studies

Low-end NetApp interconnect events exceed disk events, 4,338 to 3,230

In Jiang et al., FAST 2008, Table 1, low-end NetApp systems recorded 4,338 physical-interconnect failure events and 3,230 disk failure events from January 2004 through August 2007 (22,031 systems and 264,983 FC disks ever installed). For Figure 4(b), which excludes subsystems that used Disk H, that class has a stated disk AFR of 0.9% and a storage-subsystem AFR of 4.6%. In Maneas et al., FAST 2020, Table 2, SCSI error is 32.78% of SSD replacements (ARR 0.055%) on almost 1.4 million NetApp SSDs; that signal is a hardware error the SSD reports (the example is drive DRAM ECC), while aborted commands are 13.56% (ARR 0.023%) and include a host or connection case where write data never reaches the device.

Download CSV33 rows · 18 columns · CSV

File /data/datasets/what-fails-comparison.csv · JSON metadata · All datasets · afr · nearline-hdd · disk-shelf · enterprise-ssd

Observatory card: nearline-hdd

Method

The CSV transcribes stated figures. It does not digitize a chart and it does not refit a rate. Jiang et al. report NetApp AutoSupport logs from about 39,000 commercially deployed storage systems over 44 months, January 2004 through August 2007, with about 1,800,000 disks in about 155,000 shelf enclosures. Table 1 prints the class counts and the four failure-event counts. Those event rows are the jiang-table1 cut. The disk column is disks ever installed during the 44 months, including replacements, not a one-day census.

Figure 4(b) excludes subsystems that used the problematic Disk H family. The AFR cells are the figures written in that section, including Finding (2): near-line disk AFR 1.9%, near-line subsystem AFR 3.4%, low-end disk AFR 0.9%, low-end subsystem AFR 4.6%. Nearby sentences also say “about” for the 1.9%, 3.4%, and 4.6% figures. The low-end example says disks are about 20% of that subsystem AFR. Those four rows are jiang-figure4b. They are not Table 1 events, and this file does not divide 0.9 by 4.6.

The share rows are the cross-class prose ranges: disk failures 20–55% of storage-subsystem failures, physical interconnects including shelf enclosures 27–68%, protocol failures 5–10%, and performance failures 4–8%. They are not percentages of the Table 1 event counts. Annual rate is AFR only on the Figure 4(b) rows. ARR is a different definition, failures divided by device-years, and appears only on the Maneas rows.

Maneas et al. use ten NetApp Active IQ snapshots (previously AutoSupport), January and June 2017, January, May, August, and December 2018, and February, March, April, and May 2019: almost 1.4 million enterprise SSDs, three manufacturers, 18 models, over 30 months. Table 2 prints nine replacement reasons. The reason was missing for 40% of replacements, and the percent of replacements is normalized for that gap. SCSI error is a hardware error the SSD reports; the paper’s example is drive DRAM ECC. Aborted commands are the host-or-connection case. There is no line-rate or throughput column in this file.

Limits

Counted here

From the Table 1 rows, not from the paper’s rounded “about” lines: 39,115 systems, 156,990 shelf enclosures, 1,819,423 disks ever installed. The paper says about 39,000 systems, about 155,000 shelf enclosures, and about 1,800,000 disks. This file does not force those sentences to equal the sums.

Low-end physical-interconnect events (4,338) exceed low-end disk events (3,230). Near-line disk events are 10,105 against 4,888 physical-interconnect events, so the low-end ordering is not the ordering in every class.

The nine Table 2 share cells sum to 100. Predictive failures, threshold exceeded, and recommended failures sum to 34.44% of replacements. The paper calls that preventative group about one third and does not print 34.44. The nine ARR cells sum to 0.166. Table 2 does not print that sum, so this file does not call 0.166% the fleet annual replacement rate.

Comparison

Vendor telemetry as stated in the papers. 33 rows, the same rows as the CSV. Failure events (count) are Jiang Table 1 only. Share (%) is either a cross-class prose range, the low-end about-20% of subsystem AFR, or a normalized percent of SSD replacements. Annual rate (%) is AFR or ARR, never both, and never a line rate. “Not stated” means the source did not give that figure for the row.
CutStudyFleetWindowClassPopulationFaultSystems (count)Shelf enclosures (count)Disks ever installed (count)Failure events (count)Share, low (%)Share, high (%)Annual rate (%)Annual-rate kindSignalWhat the signal missesCitation
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08near-lineTable 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.disk failure4,92733,681520,77610,105not statednot statednot statednot statedRAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors.Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08near-lineTable 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.physical interconnect failure4,92733,681520,7764,888not statednot statednot statednot statedDisks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error.An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08near-lineTable 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.protocol failure4,92733,681520,7761,819not statednot statednot statednot statedDisks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers.The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08near-lineTable 1, 44 months (January 2004 through August 2007). Media: SATA. Multipathing: single path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.performance failure4,92733,681520,7761,080not statednot statednot statednot statedThe storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping.The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08low-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.disk failure22,03137,260264,9833,230not statednot statednot statednot statedRAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors.Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08low-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.physical interconnect failure22,03137,260264,9834,338not statednot statednot statednot statedDisks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error.An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08low-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.protocol failure22,03137,260264,9831,021not statednot statednot statednot statedDisks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers.The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08low-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.performance failure22,03137,260264,9831,235not statednot statednot statednot statedThe storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping.The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08mid-rangeTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.disk failure7,15452,621578,9808,989not statednot statednot statednot statedRAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors.Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08mid-rangeTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.physical interconnect failure7,15452,621578,9807,949not statednot statednot statednot statedDisks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error.An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08mid-rangeTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.protocol failure7,15452,621578,9802,298not statednot statednot statednot statedDisks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers.The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08mid-rangeTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.performance failure7,15452,621578,9802,060not statednot statednot statednot statedThe storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping.The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08high-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.disk failure5,00333,428454,6848,240not statednot statednot statednot statedRAID-layer disk-failure event. Includes a proactive fail when on-disk health statistics show too many sector errors.Not every fault reaches the RAID layer. The paper's example is an interconnect failure recovered by SCSI retries or tolerated by multipathing.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08high-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.physical interconnect failure5,00333,428454,6847,395not statednot statednot statednot statedDisks appear missing. The RAID layer logs a disk-missing event. Stated causes: host-adapter failure, a broken cable, shelf power loss, a shelf backplane error, or a shelf FC-driver error.An interconnect failure recovered through SCSI-layer retries, or tolerated through multipathing, is not counted as a storage-subsystem failure.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08high-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.protocol failure5,00333,428454,6841,576not statednot statednot statednot statedDisks stay visible, but I/O is not correctly answered. Stated causes: protocol incompatibility, or a software bug in the disk drivers.The paper states this tag from that symptom. Faults recovered before the RAID layer are outside the subsystem-failure count.jiang-fast08
jiang-table1jiang-fast08NetApp AutoSupport2004-01/2007-08high-endTable 1, 44 months (January 2004 through August 2007). Media: FC. Multipathing: single-path and dual-path. RAID-4 and RAID-6. The table note does not say Disk H was removed. The disk column is disks ever installed, including replacements, not a one-day census.performance failure5,00333,428454,684153not statednot statednot statednot statedThe storage layer sees a disk that does not serve I/O in time, and none of the other three failure types is detected. Stated causes: a partial failure, unstable connectivity, or disk-level recovery such as broken-sector remapping.The tag is used only when the other three types were not detected. Faults recovered before the RAID layer are outside the count.jiang-fast08
jiang-figure4bjiang-fast08NetApp AutoSupport2004-01/2007-08near-lineFigure 4(b) prose for near-line systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count.disk failurenot statednot statednot statednot statednot statednot stated1.9AFRFinding (2) states a 1.9% disk AFR. The next paragraph, still on Figure 4(b), says near-line systems mostly use SATA disks and experience about 1.9% AFR for disks.This row excludes subsystems that used Disk H. It is not the Table 1 event count. The prose does not give a near-line disk share of subsystem AFR.jiang-fast08
jiang-figure4bjiang-fast08NetApp AutoSupport2004-01/2007-08near-lineFigure 4(b) prose for near-line systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count.storage subsystemnot statednot statednot statednot statednot statednot stated3.4AFRFinding (2) states a 3.4% storage-subsystem AFR. A later sentence on the same figure says the near-line storage-subsystem AFR is about 3.4%, lower than the low-end figure.Disk H is excluded. The paper does not print this AFR as a sum of Table 1 events.jiang-fast08
jiang-figure4bjiang-fast08NetApp AutoSupport2004-01/2007-08low-endFigure 4(b) prose for low-end systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count.disk failurenot statednot statednot statednot stated20200.9AFRThe Figure 4(b) example says the disk AFR is only 0.9%, about 20% of overall subsystem AFR. Finding (2) states 0.9%. A later sentence says disk AFR for low-end, mid-range, and high-end together is under 0.9%.The 20% cell is the paper's about-20% of subsystem AFR on this Disk H-excluded cut, not a share of Table 1 events. This file does not give mid-range or high-end a point disk AFR.jiang-fast08
jiang-figure4bjiang-fast08NetApp AutoSupport2004-01/2007-08low-endFigure 4(b) prose for low-end systems, excluding storage subsystems that used Disk H. Stated in prose, not read off the chart, and not a Table 1 event count.storage subsystemnot statednot statednot statednot statednot statednot stated4.6AFRThe Figure 4(b) example says the storage-subsystem AFR is about 4.6%. Finding (2) states 4.6%, higher than the near-line subsystem AFR.Disk H is excluded. Not a Table 1 event sum. The 0.9% disk AFR is a different row.jiang-fast08
jiang-sharejiang-fast08NetApp AutoSupport2004-01/2007-08all four classesProse across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4.disk failurenot statednot statednot statednot stated2055not statednot statedDisk failures contribute to 20-55% of storage subsystem failures.Cross-class range in prose. Not the low-end about-20% point, and not a percentage of Table 1 events.jiang-fast08
jiang-sharejiang-fast08NetApp AutoSupport2004-01/2007-08all four classesProse across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4.physical interconnect failurenot statednot statednot statednot stated2768not statednot statedPhysical interconnects, including shelf enclosures, account for 27-68% of storage subsystem failures. The Figure 4(b) discussion says that fraction ranges from 27% to 68%.Cross-class prose range. Not a Table 1 event share. An interconnect fault recovered by SCSI retries or tolerated by multipathing is not in the subsystem count.jiang-fast08
jiang-sharejiang-fast08NetApp AutoSupport2004-01/2007-08all four classesProse across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4.protocol failurenot statednot statednot statednot stated510not statednot statedProtocol failures contribute to 5-10% of storage subsystem failures.Cross-class prose range. The tag means disks are visible but I/O is not correctly answered.jiang-fast08
jiang-sharejiang-fast08NetApp AutoSupport2004-01/2007-08all four classesProse across near-line, low-end, mid-range, and high-end. Not a per-class share of Table 1 events, and not a digitization of Figure 4.performance failurenot statednot statednot statednot stated48not statednot statedPerformance failures contribute to 4-8% of storage subsystem failures.Cross-class prose range. Counted when the disk does not serve I/O in time and the other three types were not detected.jiang-fast08
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.SCSI errornot statednot statednot statednot stated32.7832.780.055ARRThe SCSI layer reports a hardware error from the SSD that is severe enough to replace the drive and reconstruct the data. The example is ECC errors from the drive's DRAM.Normalized because the reason was missing for 40% of replacements. This is not the aborted-command row. The paper's host or connection case, where write data never reaches the device, is that other row.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.unresponsive drivenot statednot statednot statednot stated0.600.600.001ARRThe drive has completely failed and become unresponsive.Normalized share, not a raw ticket count. The prose says this reason is reported for only 0.60% of replacements.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.lost writesnot statednot statednot statednot stated13.5413.540.023ARRA 4K WAFL block read from the SSD is inconsistent with its signature, which includes attributes and a version number. The replace heuristic is multiple such errors on one SSD and none on any other SSD.The paper says the root cause could be a firmware bug in the drive or another layer in the stack, and that it is less clear whether the drive or other layers are to blame.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.aborted commandsnot statednot statednot statednot stated13.5613.560.023ARRAn aborted command reported either by the SSD or by the storage layer.The stated case is a host or connection issue: the host sends write commands, but the data never reach the device. The tag does not by itself prove an SSD hardware fault.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.disk-ownership I/O errorsnot statednot statednot statednot stated3.273.270.005ARRAn error while communicating with the subsystem that tracks which node owns a disk. The SSD is then marked failed.Normalized share. The paper ties the tag to that ownership subsystem. It does not give a NAND or DRAM example for this row.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.command timeoutsnot statednot statednot statednot stated1.811.810.003ARRThe SSD's own timers, and the storage layer's timers, expire. The operation does not finish in the allotted time even after retries.Normalized share. The definition does not identify the timeout as media wear-out.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.predictive failuresnot statednot statednot statednot stated12.7812.780.021ARRThe SSD reports a pattern of recovered errors using the manufacturer's own thresholds and criteria.Preventative: the prose says category D drives were still operational before replacement. A further split of predictive replacements is not in Table 2. The paper says the most common trigger behind a preventative replacement is exceeding the threshold of consecutive timeouts.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.threshold exceedednot statednot statednot statednot stated12.7312.730.020ARRThe Storage Health Monitor replaces the SSD when a threshold is crossed. The example is the number of media errors.Preventative. The drive was still operational. Recommended failures are described as less strict and less urgent than this row.maneas-fast20
maneas-table2maneas-fast20NetApp Active IQ2017-01/2019-05ssd sampleTable 2. Almost 1.4 million NetApp enterprise SSDs, three manufacturers and 18 models, over 30 months. Ten Active IQ snapshots: Jan/Jun 2017, Jan/May/Aug/Dec 2018, and Feb/Mar/April/May 2019. The reason was missing for 40% of replacements, so each percent of replacements is normalized. The table prints no raw replacement count.recommended failuresnot statednot statednot statednot stated8.938.930.015ARRThe system reports that the drive should be replaced in the near future.Less strict and less urgent than threshold exceeded. The drive was still operational before replacement.maneas-fast20

Window, fleet, and population are columns in the CSV. Jiang rows use NetApp AutoSupport, 2004-01/2007-08. Maneas rows use NetApp Active IQ, 2017-01/2019-05. The population cell states which cut the numbers belong to.

Columns

row_kind (string)
Which stated cut the row transcribes: jiang-table1, jiang-figure4b, jiang-share, or maneas-table2.
study (string)
Citation id of the paper that states the row.
fleet (string)
Named fleet: NetApp AutoSupport or NetApp Active IQ.
window (string)
Study window as year-month bounds. Jiang is January 2004 through August 2007 (44 months). Maneas is the 30 months covered by 10 snapshots from January 2017 through May 2019.
system_class (string)
Jiang storage-system class (near-line, low-end, mid-range, high-end, or all four classes) or ssd sample.
population (string)
Who was counted, which cut (Table 1, Figure 4(b) excluding Disk H, cross-class prose, or Table 2), and what the count is not.
fault (string)
Named failure or replacement reason, using the paper's categories.
systems_count (integer)
Systems (count) from Jiang Table 1. Empty when the row is not a Table 1 class row. Not a Figure 4(b) population.
shelves_count (integer)
Shelf enclosures (count) from Jiang Table 1. Empty when the row is not a Table 1 class row.
disks_ever_installed_count (integer)
Disks ever installed (count) during the 44 months, including replacements, from Jiang Table 1. Empty when the row is not a Table 1 class row. Not the Maneas SSD population.
failure_events (integer)
Failure events (count) for that fault in Jiang Table 1. Empty when the source did not state an event count for the row.
share_low_percent (number)
Share, low (%). Jiang cross-class prose uses a range. The low-end Figure 4(b) disk row uses 20 for the paper's about-20% of subsystem AFR. Maneas Table 2 percent of replacements is a point, so low equals high. Empty when not stated.
share_high_percent (number)
Share, high (%). Equal to share_low_percent when the source stated a point rather than a range. Empty when not stated.
annual_rate_percent (number)
Annual rate (%). AFR for Jiang Figure 4(b) prose, or ARR for Maneas Table 2. Empty when the source did not state an annual rate for the row. Not a line rate and not a throughput.
annual_rate_kind (string)
AFR or ARR when annual_rate_percent is set. Empty otherwise. AFR and ARR are different definitions and are not converted into each other.
signal (string)
The log or telemetry signal the paper says catches this fault.
misses (string)
What that signal does not prove, from a sentence in the same paper.
citation_id (string)
Id into the citations list. Open that URL for the sentence.

License

Small derived table of figures stated in Jiang et al., FAST 2008, and Maneas et al., FAST 2020. Not a Creative Commons license. The opened FAST '14 Consent Form for Refereed Papers says USENIX lets authors retain copyright and asks only for a right to publish. That grant is non-exclusive, worldwide, perpetual, and irrevocable, and also non-sublicensable and non-transferable. The form contains no CC grant. This CSV is not the NetApp AutoSupport database and not the Active IQ bundles. Neither paper links a public dump or a data DOI for those logs. Canonical files are the USENIX PDFs.

FAST ’14 refereed-paper consent form · Canonical FAST 2008 PDF (AutoSupport logs are not in the file) · Canonical FAST 2020 PDF (Active IQ bundles are not in the file)

Sources

  1. Are Disks the Dominant Contributor for Storage Failures? A Comprehensive Study of Storage Subsystem Failure Characteristics (accessed 2026-10-08)
  2. A Study of SSD Reliability in Large Scale Enterprise Storage Deployments (accessed 2026-10-08)
  3. FAST '14 Consent Form for Refereed Papers (accessed 2026-10-08)