Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Dataset · public postmortems · SoCC 2016

355 of 597 public cloud outages have an unknown root cause

Gunawi et al., SoCC 2016, Table 3, mark 355 of 597 unplanned outages as UNKNOWN. The study covers 32 popular Internet services and 1,247 public reports from January 2009 through December 2015. The paper states that only 40% of those descriptions reveal a root cause and only 24% reveal a fix. Among known root-cause tags, the three largest counts are UPGRADE 54 (16%), NETWORK 52 (15%), and BUGS 51 (15%). STORAGE is 13 occurrences on 4 services (4% of known causes). HARDWARE is 5 occurrences on 4 services (1% of known causes). One outage can carry more than one tag, so the occurrence column is not a partition of the 597 outages. In the yearly column, M means at least 10, not exactly 10.

Download CSV39 rows · 16 columns · CSV

File /data/datasets/postmortem-unknown-causes.csv · JSON metadata · All datasets · what fails in storage subsystems

Method

The CSV transcribes stated figures from Gunawi et al., SoCC 2016. It does not digitize a chart and it does not refit a percent. The population is the paper's cloud outage study: 32 Internet services named in Table 1, searched as "serviceName outage month year" on Google and Bing, first 30 hits, every month from January 2009 through December 2015. That search produced 1,247 unique links describing 597 unplanned outages. Each outage was read by at least four authors. A vague "configuration problem" is tagged CONFIG only. The authors say they do not speculate.

Table 3 is the table3 block. Occurrences (count) are the Cnt column. Services with this root cause (count) are the #Sv column. Share (%) is the paper's percent among known root causes. The caption says the percent is not computed on the UNKNOWN row, so that share cell is empty. One outage can carry more than one root-cause tag, so the occurrence column is not a split of the 597 outages. Yearly marks are the seven-token string for 2009 through 2015, copied as printed. M means at least 10. A dash stays a dash.

The uptime rows are section 3.2 prose, not a reading of Figure 1 or Figure 2. The paper's 99% line is not more than 88 hours of annual downtime. The 99.9% line is not more than 8.8 hours. Those ceilings are definitions, not measured downtimes, and they stay in the signal text rather than in a count column. Impact rows and fix rows are the percents printed in section 6. They are not Figure 5 cell counts. These figures are the paper's counts of public reports, not a provider's advertised uptime.

Limits

Counted here

From the Table 3 cells, not from a total the paper prints: the 13 known-cause occurrence counts sum to 340. Adding the 355 UNKNOWN occurrences gives 695 tag occurrences. That is not 597. The paper says an outage can be tagged with more than one root cause. 597 minus 355 is 242. That complement is not a printed cell. The paper states the disclosure rate as 40%, and this file does not replace 40 with a ratio of those two counts.

Where a yearly string contains no M, reading each dash as zero makes the seven numbers sum to the occurrence count. That check covers LOAD, POWER, SECURITY, HUMAN, STORAGE, SERVER, NATDIS, and HARDWARE. It is a check, not a second measurement, and the CSV still stores the dash. Where M appears, the hidden year is not filled in.

Comparison

Gunawi et al., SoCC 2016. Population: 32 Internet services, 597 unplanned outages, 1,247 public reports, January 2009 through December 2015. 39 rows, the same rows as the CSV. Occurrences (count) are Table 3 tags, not a partition of the 597 outages. Share (%) uses the basis column. Yearly marks are copied as printed. M means at least 10 occurrences that year, not exactly 10. A dash is the paper's dash. "Not stated" means the source has no such cell for the row.
CutMeasureSectionOutages (count)Reports (count)Metadata tags (count)Services in the study (count)Services with this cause (count)Services short of uptime (count)Occurrences (count)Share (%)Share basisYearly marks, 2009–2015 (as printed)SignalWhat the signal missesCitation
populationservices in the study2not statednot statednot stated32not statednot statednot statednot statednot statednot statedTable 1 names 32 services in chat, e-commerce, email, game, PaaS/IaaS, SaaS, social, storage, and video. When the paper prints uptime numbers it replaces names with labels such as CH2.Not an industry census. Popularity skews which outages become headlines. This file does not map the anonymized labels back to names.gunawi-socc16
populationunplanned outages1.1597not statednot statednot statednot statednot statednot statednot statednot statednot statedThe abstract and section 1.1 count 597 unplanned outages. Section 2 defines a service outage as an unplanned unavailability of full or partial features that affects all or a significant number of users and is reported publicly.Planned maintenance is excluded. Failures of non-essential operations are excluded. Unreported outages are missing, so the paper calls the availability numbers minimums.gunawi-socc16
populationpublic reports2not stated1,247not statednot statednot statednot statednot statednot statednot statednot statedSection 2: the query was serviceName outage month year on Google and Bing, for every month from January 2009 through December 2015. The authors read the first 30 hits and kept 1247 unique links. They read 1 to 12 sources per outage, 2 on average.The links are not in this file. Search rank is not a sample of all incidents. Reference [1] is not a copy of the links.gunawi-socc16
populationmetadata tags2not statednot stated3,249not statednot statednot statednot statednot statednot statednot statedSection 2: COSDB contains 597 outage descriptions, 1247 links, and 3249 outage metadata tags. The paper says COSDB is downloadable from reference [1], http://ucare.cs.uchicago.edu/projects/cbs/.This file is not COSDB. A fetch of reference [1] on 2026-10-09 did not return a verified page, because the TLS certificate could not be validated. The live contents of that URL are not described here.gunawi-socc16
populationdescriptions that reveal a root cause7not statednot statednot statednot statednot statednot statednot stated40outage descriptions, as stated in section 7not statedSection 7 states that only 40% of outage descriptions reveal root causes. Section 5 states that 355 of 597 outages have UNKNOWN root causes.40% is the paper's statement, not a recomputed ratio. The complement 597 minus 355 is not a printed cell. This percent is of descriptions, not of known-cause tags.gunawi-socc16
populationdescriptions that reveal a fix procedure6not statednot statednot statednot statednot statednot statednot stated24outage descriptions, as stated in section 6not statedSection 6 states that only 24% of outage descriptions reveal the fix procedures. Section 7 repeats that sentence.The eight fix rows are shares of reported procedures, not of all descriptions. 24% is not a repair success rate.gunawi-socc16
populationoutages reported with downtime3.2not statednot statednot statednot statednot statednot statednot stated69the 597 outagesnot statedSection 3.2, and also section 1.1, states that 69% of outages are reported with downtime information. The uptime figures use that subset. The sentence says within this population.Duration is missing for the rest. Figures 1, 2, and 4 are not digitized. This study states no line rate and no throughput.gunawi-socc16
uptimeworst year short of 99 percent uptime3.2not statednot statednot statednot statednot stated10not stated31the 32 services, worst year from 2009 through 2015not statedSection 3.2, worst year of each service: 10 services (31%) do not reach two-nine uptime. The paper's 99% line is not more than 88 hours of annual downtime.A minimum, because not all outages are public and not all public reports include downtime. Not the average-year figure. Service names stay anonymized. Not five-nine uptime, which the paper defines as five minutes of annual downtime and calls still far from reach.gunawi-socc16
uptimeworst year short of 99.9 percent uptime3.2not statednot statednot statednot statednot stated27not stated84the 32 services, worst year from 2009 through 2015not statedThe same worst-year cut: 27 services (84%) do not reach three-nine uptime. The paper's 99.9% line is not more than 8.8 hours of annual downtime.Same reporting minimum as the 99% worst-year row. Not a named service's downtime. Not the average-year count of 25 services.gunawi-socc16
uptimeaverage year short of 99 percent uptime3.2not statednot statednot statednot statednot stated2not stated6the 32 services, the paper's average across the six yearsnot statedSection 3.2 prints a second cut: on average across the six years, 2 services (6%) do not reach 99% uptime, using the ceiling of 88 hours of annual downtime.The search window is seven years, January 2009 through December 2015. This file does not change the paper's phrase six years. Not the worst-year count of 10 services.gunawi-socc16
uptimeaverage year short of 99.9 percent uptime3.2not statednot statednot statednot statednot stated25not stated78the 32 services, the paper's average across the six yearsnot statedThe same sentence: 25 services (78%) do not reach 99.9% uptime on that average, using the ceiling of 8.8 hours of annual downtime.Same six-year wording as printed. Not the worst-year count of 27 services.gunawi-socc16
table3UNKNOWNnot statednot statednot statednot statednot stated29not stated355not statednot statedM.M.M.M.M.M.MTable 3 Cnt is 355 outages out of 597 with UNKNOWN root causes, across 29 of the 32 services. The yearly marks are M.M.M.M.M.M.M. The caption says M means at least 10 in that year.The tag is the absence of a concrete root cause in the public report. It does not show that operators never found a cause. Seven exact tens would be 70, not 355, so the marks are not yearly counts. The authors say they do not speculate past the report.gunawi-socc16
table3UPGRADE5.1not statednot statednot statednot stated18not stated5416known root-cause tags7.4.M.5.M.4.7Section 5.1 uses UPGRADE for hardware upgrades or software updates, typically during maintenance. Table 3 prints 18 services, 54 occurrences, and 16% of known causes, the largest known count.The tag does not say whether the new software, the upgrade script, or a later load failed. The same outage can also carry BUGS, CONFIG, HUMAN, NETWORK, or LOAD. Each M is at least 10, not an exact year. Not a count of all maintenance windows.gunawi-socc16
table3NETWORK5.2not statednot statednot statednot stated21not stated5215known root-cause tags4.4.6.8.M.8.5Section 5.2 covers networking failures, including a dead core switch, redundant paths that fail together, access misconfiguration, failed upgrades, and DNS, traffic-control, or routing faults when the report describes a network problem. Table 3 prints 21 services, 52 occurrences, and 15%.External networks outside the service are included. A network bug is not retagged BUGS unless the report says bug or software error. Not a packet-loss or throughput measurement. The M mark is at least 10, not an exact year.gunawi-socc16
table3BUGS5.3not statednot statednot statednot stated18not stated5115known root-cause tagsM.4.9.8.9.9.2Section 5.3 tags BUGS only when the report explicitly says bugs or software errors. Table 3 prints 18 services, 51 occurrences, and 15% of known causes.The paper says many other cases can be traced to software but are not labeled BUGS unless the report says so, so the ratio can be larger than 15%. The M mark is at least 10, not an exact 10.gunawi-socc16
table3CONFIG5.4not statednot statednot statednot stated19not stated3410known root-cause tags2.2.7.2.5.M.4Section 5.4, with the method rule in section 2: a report that says a configuration problem is tagged CONFIG and not also BUGS or HUMAN. Table 3 prints 19 services, 34 occurrences, and 10%.The tag does not separate a bug that corrupted configuration from an operator who set the wrong value. The authors say they do not speculate. The M mark is at least 10.gunawi-socc16
table3LOAD5.5not statednot statednot statednot stated18not stated319known root-cause tags2.5.5.5.4.8.2Section 5.5: unexpected traffic overload from user requests, code upgrades, misconfiguration, or flawed recovery that adds traffic, including a recovery storm. Table 3 prints 18 services, 31 occurrences, and 9%.Not a measured request rate and not a capacity. Recovery code that behaves as designed can still be in this tag when it creates extra traffic. No M appears. The seven yearly numbers sum to 31.gunawi-socc16
table3CROSS5.6not statednot statednot statednot stated14not stated288known root-cause tags-.2.4.M.5.3.4Section 5.6: a disruption from another service, such as an ISP, a lower layer, or a third party. Table 3 prints 14 services, 28 occurrences, and 8%.This is not the internal root cause of the affected service. The fix tag NOTHING is a different cut. The M mark is at least 10.gunawi-socc16
table3POWER5.7not statednot statednot statednot stated11not stated216known root-cause tags5.4.3.5.3.1.-Section 5.7: power failures, including lightning storms, a vehicle hitting utility poles, failed utility maintenance, and backup generators that do not take the load. Table 3 prints 11 services, 21 occurrences, and 6%.A second failure of the backup path stays in this tag. Not a utility outage census. The paper says POWER tends to produce long downtimes because backups are often affected. That duration claim is not a Figure 4 digit. The positive yearly digits sum to 21 if the final dash is zero.gunawi-socc16
table3SECURITY5.8not statednot statednot statednot stated9not stated175known root-cause tags7.-.2.1.3.4.-Section 5.8: security-related causes, including DDoS, a worm, and a botnet, when the report presents them as the cause. Table 3 prints 9 services, 17 occurrences, and 5%.Not the impact category SECURITY, which section 6 prints as 1% of impacts. An example attack size is not a column in this file. The positive yearly digits sum to 17 if each dash is zero.gunawi-socc16
table3HUMAN5.9not statednot statednot statednot stated11not stated144known root-cause tags-.1.4.4.2.1.2Section 5.9: a human error on a manual path. Table 3 prints 11 services, 14 occurrences, and 4%.The method does not add HUMAN when the report only says configuration problem. The section says more than half of human errors relate to upgrade and configuration. That half is prose, not a Table 3 cell, so this file does not store it. The positive yearly digits sum to 14 if the leading dash is zero.gunawi-socc16
table3STORAGE5.10not statednot statednot statednot stated4not stated134known root-cause tags2.-.-.3.5.3.-Section 5.10 tags STORAGE only when the report explicitly mentions a failure from the storage layer: devices, the storage cluster, or a file or database system. Table 3 prints 4 services, 13 occurrences, and 4% of known causes.A storage fault described only as a server, a generic hardware failure, or a bug is not this tag. The 4% is among known root-cause tags, not among the 597 outages. Not a disk AFR and not a drive count. The paper does not define the dash. The positive digits 2, 3, 5, and 3 sum to 13.gunawi-socc16
table3SERVER5.11not statednot statednot statednot stated6not stated113known root-cause tags-.3.-.2.2.4.-Section 5.11 labels SERVER when the report mentions a node or server but not a specific component failure. The section's examples are caching, LDAP, login, quota-check, scheduling, and search nodes. Table 3 prints 6 services, 11 occurrences, and 3%.Specific component failures are outside this tag. The section number is shared with HARDWARE under Miscellaneous Server and Hardware Failures. The positive digits 3, 2, 2, and 4 sum to 11.gunawi-socc16
table3NATDIS5.12not statednot statednot statednot stated5not stated93known root-cause tags1.1.3.2.1.1.-Section 5.12: external and natural disasters. The section names lightning, a vehicle crashing into utility poles, construction cutting optical cables, and an undersea fiber cut. Table 3 prints the label as NatDis, 5 services, 9 occurrences, and 3%.The same lightning event can also be a POWER tag. Not a weather dataset. The paper says NATDIS tends toward long downtimes. Figure 4 is not digitized. The positive digits sum to 9. Section 5.12 follows the shared section 5.11.gunawi-socc16
table3HARDWARE5.11not statednot statednot statednot stated4not stated51known root-cause tags1.-.-.3.1.-.-Section 5.11 labels HARDWARE when the report does not name a specific type of hardware failure. Table 3 prints the same section number 5.11 as SERVER. 4 services, 5 occurrences, and 1% of known causes.In most of these cases the providers do not present the details. A named switch, disk, or power fault is NETWORK, STORAGE, or POWER instead. The repeated 5.11 is the paper's table, not a section 5.13. The population is small. The paper exempts HARDWARE from the claim that almost every root cause has a maximum downtime above 50 hours. Figure 4 is not digitized. The positive digits 1, 3, and 1 sum to 5.gunawi-socc16
impactFULLOUTAGE6not statednot statednot statednot statednot statednot statednot stated59the six impact categories in section 6not statedSection 6 prints full outages as 59% of the impact breakdown. Section 2 defines FULLOUTAGE as a full service outage.The six printed impact percents sum to 99, not 100. The paper does not say the denominator is the 597 outages. Figure 5 counts are not transcribed. Not a root-cause tag.gunawi-socc16
impactOPFAIL6not statednot statednot statednot statednot statednot statednot stated22the six impact categories in section 6not statedSection 6 prints failures of essential operations as 22%. Section 2's examples of essential operations are login, payment, and search. Non-essential examples, such as a profile-picture update, are excluded from the outage set.Same impact-percent limit as FULLOUTAGE. Not a count of failed requests.gunawi-socc16
impactPERFORMANCE6not statednot statednot statednot statednot statednot statednot stated14the six impact categories in section 6not statedSection 6 prints performance glitches as 14%. Section 2 treats late deliveries that lead to loss of productivity as an outage, labeled PERFORMANCE.Not a latency distribution and not a throughput. The six impact shares sum to 99 as printed.gunawi-socc16
impactLOSS6not statednot statednot statednot statednot statednot statednot stated2the six impact categories in section 6not statedSection 6 prints data loss as 2%. Section 2 lists LOSS as an impact tag.Not a count of bytes lost. Not the root-cause tag STORAGE.gunawi-socc16
impactSTALE6not statednot statednot statednot statednot statednot statednot stated1the six impact categories in section 6not statedSection 6 prints data staleness or inconsistency as 1%. Section 2 lists STALE as an impact tag.Not a consistency-model measurement. Same impact-breakdown limit as the other five impact rows.gunawi-socc16
impactSECURITY6not statednot statednot statednot statednot statednot statednot stated1the six impact categories in section 6not statedSection 6 prints security attacks or breaches as 1% of the impact breakdown.Not the root-cause tag SECURITY, which Table 3 prints as 17 occurrences and 5% of known causes. The two uses of the word are different cuts.gunawi-socc16
fixADDRESOURCES6not statednot statednot statednot statednot statednot statednot stated10reported fix procedures in section 6not statedSection 6 prints add additional resources as 10% of reported fix procedures. Table 2's tag is ADDRESOURCES.Only 24% of descriptions reveal a fix, so this percent is not a share of the 597 outages. The eight printed fix percents sum to 99. Not a capacity that was added. Figure 5b counts are not transcribed.gunawi-socc16
fixFIXHW6not statednot statednot statednot statednot statednot statednot stated22reported fix procedures in section 6not statedSection 6 prints fix hardware as 22% of reported fix procedures. Table 2's tag is FIXHW.Not the root-cause tag HARDWARE, which is 5 occurrences. A hardware fix does not mean the root cause was HARDWARE. Among reported fixes only.gunawi-socc16
fixFIXSW6not statednot statednot statednot statednot statednot statednot stated22reported fix procedures in section 6not statedSection 6 prints fix software as 22% of reported fix procedures. Table 2's tag is FIXSW.Not the root-cause tag BUGS. A software fix does not mean the report said bug. Among reported fixes only.gunawi-socc16
fixFIXCONFIG6not statednot statednot statednot statednot statednot statednot stated7reported fix procedures in section 6not statedSection 6 prints fix misconfiguration as 7% of reported fix procedures. Table 2's tag is FIXCONFIG.Not the root-cause tag CONFIG, which is 34 occurrences and 10% of known causes. Among reported fixes only.gunawi-socc16
fixRESTART6not statednot statednot statednot statednot statednot statednot stated4reported fix procedures in section 6not statedSection 6 prints restart of affected components as 4% of reported fix procedures. Table 2's tag is RESTART.A restart does not name the failed component. Among reported fixes only. This file does not give a restart procedure.gunawi-socc16
fixRESTOREDATA6not statednot statednot statednot statednot statednot statednot stated14reported fix procedures in section 6not statedSection 6 prints restore data as 14% of reported fix procedures. Table 2's tag is RESTOREDATA.Not a count of bytes restored. Among reported fixes only. This file does not describe a restore procedure.gunawi-socc16
fixROLLBACKSW6not statednot statednot statednot statednot statednot statednot stated8reported fix procedures in section 6not statedSection 6 prints rollback software as 8% of reported fix procedures. Table 2's tag is ROLLBACKSW.Not the root-cause tag UPGRADE. A rollback does not by itself prove an upgrade was the cause. Among reported fixes only.gunawi-socc16
fixNOTHING6not statednot statednot statednot statednot statednot statednot stated12reported fix procedures in section 6not statedSection 6 prints nothing, due to cross-dependencies, as 12% of reported fix procedures. Table 2's tag is NOTHING.Not the root-cause tag CROSS, which is 28 occurrences and 8% of known causes. Nothing means the reported fix was blocked by another service, not that no harm occurred. Among reported fixes only.gunawi-socc16

HARDWARE and SERVER both use section 5.11 because both sit under "Miscellaneous Server and Hardware Failures." NATDIS is section 5.12. The Jiang et al. comparison of disk and interconnect events is a separate dataset: what-fails-comparison.

Columns

row_kind (string)
Which stated cut the row transcribes: population, uptime, table3, impact, or fix.
measure (string)
The paper's name for the row: a study total, an uptime cut, a Table 3 root-cause tag, an impact tag, or a fix tag.
section (string)
Section number as printed. Empty only for UNKNOWN, which has a blank section cell in Table 3. HARDWARE and SERVER are both 5.11.
outages_count (integer)
Unplanned outages (count). Filled only on the population row for the 597 outages. Empty means the paper does not give this row as an outage total.
reports_count (integer)
Public reports (count). Filled only on the population row for the 1247 unique links. Empty otherwise.
tags_count (integer)
Outage metadata tags (count). Filled only on the population row for the 3249 COSDB tags. Empty otherwise. Not a sum of Table 3.
study_services_count (integer)
Services in the study (count). Filled only on the population row for the 32 services in Table 1. Empty otherwise.
tagged_services_count (integer)
Services with this root cause (count). Table 3 column #Sv. Empty when the row is not a Table 3 root cause.
services_short_of_uptime_count (integer)
Services short of the stated uptime (count), from section 3.2. Empty when the row is not an uptime cut.
occurrences_count (integer)
Root-cause occurrences (count). Table 3 column Cnt. Not a partition of the 597 outages. Empty outside Table 3.
share_percent (integer)
Printed share (%). The basis is share_basis, not a single denominator. Empty when the paper prints a dash or does not print a percent. Table 3 UNKNOWN is empty.
share_basis (string)
What share_percent is a percent of, using the paper's wording. Empty when share_percent is empty.
yearly_marks (string)
Table 3 yearly marks for 2009 through 2015, seven tokens separated by periods, copied as printed. M means at least 10, not exactly 10. A dash is the paper's dash, not a zero stored by this file. Empty outside Table 3.
signal (string)
The report wording or column the paper uses to assign this tag or count.
misses (string)
What that signal does not prove, from a sentence in the same paper or from a blank the paper leaves.
citation_id (string)
Id into the citations list. Open that URL for the sentence. Every row cites the SoCC 2016 PDF.

License

Small derived table of figures stated in Gunawi et al., SoCC '16, with credit. Not a Creative Commons license. The first page of the PDF that was opened says abstracting with credit is permitted. It also says that to copy otherwise, or to republish, post on servers, or redistribute to lists, requires prior specific permission and/or a fee, requested from permissions@acm.org. Copyright is held by the owner/author(s). Publication rights are licensed to ACM. DOI as printed: 10.1145/2987550.2987583. This file is not the paper, not the 1247 reports, and not COSDB. The paper says COSDB is downloadable from reference [1], http://ucare.cs.uchicago.edu/projects/cbs/. A fetch of that URL on 2026-10-09 did not return a verified page.

Canonical SoCC 2016 PDF used for this table · DOI as printed on that PDF: 10.1145/2987550.2987583

Sources

  1. Why Does the Cloud Stop Computing? Lessons from Hundreds of Service Outages (accessed 2026-10-09)
  2. Paper Review: Why Does the Cloud Stop Computing? Lessons from Hundreds of Service Outages (accessed 2026-10-09)