Ranking of outage types: what harms the quality of infrastructure the most
To the on-call team, the main source of trouble is what breaks most often: hardware, network links. But if you look in the log not at the number of cases but at lost hours, the picture is different. These are two different rankings, and in almost every company where we counted them, they do not match. What harms the infrastructure most is rarely first by number of events.
Below are twelve sources of outages ranked by impact. First - how you can compare such different things as a failed power supply and an instruction nobody followed. Then the ranking itself as a table, and what it rests on: our operations experience, checked against open research on outages. Then more detail on the rows that cause the most downtime. At the end - how to build such a ranking from your own data, not from an "average across the industry".
1. How to compare outages
For a single incident, ITIL has a common pair of attributes: impact (how many users and services are affected) and urgency (how fast it must be restored). Together they give the priority: which outage to handle first. To compare whole types of outages over a year, this is not enough. We use four questions. This is not an industry standard but a practical scheme built from common measures.
- How many hours of downtime added up. Frequency times average duration. The main question: it shows where the hours really go. Often a rare type of outage takes more time than a daily small one.
- How wide it hits. The scale of impact: how many services and users one failure affects. In SRE practice (site reliability engineering, Google's approach) this is called the "blast radius". A failure of one server and a power-off of a whole site can last the same time but cost very differently.
- Can you prevent it yourself. Your own changes are fully in your hands. A supplier - partly, and only through the contract. Weather and building work outside - not at all.
- Who fixes it. If the provider restores the service, your skill hardly affects the duration. Two things matter: how fast you understood that the cause is on their side, and what the contract says. In ITIL this is the area of supplier management and service level agreements.
A fifth question is useful too: could you see it in advance. A disk that was filling with logs for a month is a gap in monitoring: you could have known. A cable cut by an excavator - no. The first is fixed by monitoring, the second only by redundancy.
2. The ranking: twelve sources
The order is for a typical company: a few dozen systems, its own site or a rented one, some services from external suppliers. Frequency and share of lost time are given in words, not percentages: exact shares differ in every company, and putting in someone else's numbers is the very mistake this article warns against.
| No. | Source of outage | Frequency | Share of lost time | Who restores |
|---|---|---|---|---|
| 1 | Errors in procedures and in how operations is organisedthe instruction was not followed, does not exist or is wrong; it took long to find who is responsible | high | high | Fully you |
| 2 | Your own changes and updatesnew releases, reconfiguration, system moves | high | high | Fully you |
| 3 | Planned work and maintenancemaintenance windows, switchovers, preventive work | medium | high | Fully you |
| 4 | External suppliersnetwork links, cloud, rented sites, payment and mail gateways | medium | high | The supplier; from you - notice fast and contact them |
| 5 | Lack of capacity and overloaddemand peaks, reporting periods, background jobs | medium | medium | Fully you |
| 6 | Software bugsmemory leaks, growing logs, hung sessions | medium | medium | You and the software vendor |
| 7 | Power and coolingbuilding power feed, uninterruptible power supplies, air conditioning | low | high when it happens | You or the site owner |
| 8 | Manual edits that bypass the change process"fixed it quickly and forgot", settings drift from the reference | medium | medium | Fully you |
| 9 | Expired itemscertificates, licences, service account passwords, domains | low | medium | Fully you |
| 10 | Hardwaredisks, power supplies, boards, optics | medium | low | You and the spare parts supplier |
| 11 | Security incidentsattacks, ransomware, stolen access | low | extremely high when it happens | You, the security team, sometimes outside experts |
| 12 | External environmentweather, building work, cable cut, a failure at the neighbours | low | low | Neither of you |
Two notes. First: rows 11 and 12 are at the bottom by their usual share of lost time, not by danger. If you rank by the worst case, security and power rise to the top three. That is a different ranking and a different task - a disaster recovery plan. Second: row 1 is not really a type of outage. It acts on all other rows: it decides how long any outage lasts. That is why it is first.
What the order rests on. Google writes in its book on service reliability (Site Reliability Engineering) that about 70% of outages are caused by changes to a live system - these are rows 2, 3 and 8. Uptime Institute (outage report for 2025): almost 40% of organisations had a major outage caused by staff error in three years, and in 85% of these cases the procedure was not followed or was wrong - row 1. Same source: about two thirds of publicly known outages came from external suppliers - cloud, telecom, colocation - row 4. And again: among serious outages of data centres themselves, the first cause is power. This does not contradict row 7: power fails rarely, but it switches everything off at once.
3. Why organisation matters more than technology
The downtime of any outage is made of four parts. Detection: how long the outage went unnoticed. Response: how long it took to find who will fix it, and when that person started. Diagnosis: how long it took to find the cause. Recovery: how long the fix took. SRE and DevOps have common measures for this - mean time to detect (MTTD) and mean time to recover (MTTR). Technology affects only the last part. The first three are monitoring, on-call duty, the way an outage is passed to the right specialist, and the known error database - that is, organisation.
Situation: in two companies the same component of the same hardware fails. In the first, downtime lasts 20 minutes: the alert came at once, the on-call engineer found a description in the known error database and applied the workaround from the instruction. In the second - 4 hours: they learned about the outage from users, the on-call engineer spent an hour and a half finding who is responsible for the system, and the right engineer was on holiday, and only the manager knew his phone number. Conclusion: the same technology, a 12-times difference - all of it is organisation.
The second and third rows - your own changes and planned work - are high for another reason. These are fully controllable sources: you decide what to change, when, and with what rollback plan. So investment here pays back fastest: you need no purchases and no talks with suppliers. To measure the result, DevOps uses the change failure rate: how many changes out of a hundred led to an outage or a rollback.
4. Planned work: the most underrated source
Planned work almost always drops out of statistics for a formal reason: it is not logged as incidents. It is not in the outage log, the agreed window is usually subtracted from the availability report, and nobody talks about it in reviews. But for the client there is no difference: the service was unavailable. The difference exists only in internal reporting.
Three ways planned work creates downtime
- The window did not close on time. The work dragged on, no rollback was prepared, and the service came back not after two hours but after six. Formally it is "planned work", in fact it is a failure that nobody counts as a failure.
- A consequence showed up later. The work went well, but two days later it turned out that backups stopped working or the second link dropped. Nobody linked it to the work, and the outage went into statistics as a separate one.
- Short windows add up. Each maintenance is short and agreed, but there are many, and over a year they add up to hours comparable to failures. Nobody counts this sum, because each window is looked at separately.
The cure is not to stop the work: postponed maintenance creates failures by itself. The cure is four rules, and you can check that each one is followed.
- Count planned windows in the overall unavailability balance. As a separate line, but in the same report. Otherwise nobody sees the windows that added up.
- Check the backup before the work starts, not during it. A classic failure: work on the main node when the backup node is already broken, and nobody knew.
- Write down in advance when to roll back. In ITIL this is the rollback plan. Not "if something goes wrong", but "if the service is not up by 04:00, we return everything as it was". A decision made at night on the spot almost always sounds like "let's try a bit more".
- Keep a period of extra attention after the work. For a day or two, any outage is by default treated as linked to the work until proven otherwise. In ITIL a similar practice is called early life support after a change. Only this way do you catch consequences that show up later.
5. External providers: someone else's recovery time
Supplier outages are high in the ranking not because of frequency. It is about the fourth question: you are not the one who fixes it. The team's skill, tools and discipline hardly affect the length of downtime at that moment. Two things matter: how fast you understood the cause is outside, and what the contract says.
Situation: the service does not work, the team spends an hour and a half looking for the cause on its own side, and only then someone checks the external link or the cloud provider's status page. That hour and a half is not the supplier's time but yours, spent looking in the wrong place. What helps: checking the service from outside, through the client's eyes, not only from inside, and a list of external dependencies that the on-call engineer goes through in the first minutes.
Contracts deserve a separate word: many companies have an illusion of protection here. The recovery time in a service level agreement is not a guarantee that the service will be up in that time. It is a condition for compensation if it is not: usually a discount on the next bill. Compensation does not bring back clients who could not use your service during those hours. That is why critical suppliers are covered by redundancy - a second link from another operator, a second site, a backup gateway - and the contract serves as an extra safety net.
6. Power: rare, but all at once
Power and cooling are in the middle of the ranking, and this place looks strange: intuition puts them either higher or lower. By share of yearly downtime, power is behind changes and planned work: there are few events. By the severity of one event, nothing compares.
The reason is the scale of impact. A power failure switches off not one system but the whole site. All redundancy inside the site is useless at that moment: the backup server is in the same rack and runs on the same power feed. So protection here is built first of all into the design of the site: two independent power feeds, backup sources, systems spread across different sites, and a recovery plan.
The second feature is familiar to everyone who has been through it: after a power cut the infrastructure does not come back by itself. Some hardware does not switch on, some comes up in the wrong order, databases need checking, applications cannot find the services they need. Real downtime depends not on how long the power was off but on how well the start-up order is practised. And it is almost never tested.
A practical rule: test uninterruptible power supplies and generators under load on a schedule, and write down the start-up order of systems after a full power cut and test it at least once a year. An untested backup and an unpractised start-up are found at the worst possible moment.
7. Why hardware is at the bottom
This is the most debated row: it goes against intuition and against coffee-room talk. Hardware breaks, and regularly. But since redundancy became the norm - disks in arrays, dual power supplies, servers in clusters, backup links - the failure of one part usually does not stop the service. The failure happens, the downtime does not.
Hardware climbs up the ranking in three cases, and each of them really belongs to row 1 - the organisation of operations.
- The backup did not work. The second power supply failed six months ago, there was an alert about it, but nobody acted on it. The first failure became the last.
- No spare parts. The failed node waits a week for a part. This is not a hardware failure but a lack of stock and of a support contract.
- The hardware is past its life. The vendor no longer supports it, failures become more frequent, replacement is postponed, and each outage is reviewed as random. This is a postponed decision, not something inevitable.
From this comes a general principle for the whole ranking: the lower the row, the more often its real share is explained by row 1. Technology rarely fails on its own - it fails where the organisation left it without a safety net.
8. How to build your own ranking
Our order is a guide, not your result. For a company with one site and old hardware it is one thing. For a company that lives fully in the cloud it is another: there suppliers rise to first place, and power disappears from the list. Your own ranking is built from the incident log for a year.
- Take the log for a full year. Not for a quarter: on a short period, seasonal peaks and rare heavy outages distort the picture.
- Give each outage one source from the list. One, even if there were several causes: the one without which the outage would not have happened. Unclear cases go to a separate group. If it is big, that is already a finding: causes are recorded badly.
- Add up duration, not the number of cases. And add planned windows as a separate line, otherwise the windows that added up, from section 4, stay invisible.
- Estimate the scale of impact. How many services or users one outage of each source affects on average. This column often swaps the top rows.
- Separate what you can influence from the rest, and look at the top of the first part. This is your work plan for the next quarter: it has only what you can change.
The most common mistake is in the second step: recording the symptom instead of the source. "Service unavailable" is not a source, it is what the user saw. The source is what the outage would not have happened without: a configuration change, a full disk, a contractor's work on the neighbouring power feed. A ranking by symptoms comes down to a couple of rows, and no decisions follow from it.
A short conclusion. The quality of infrastructure is decided first of all not by what you have installed, but by what happens around failures: how fast they are seen, who takes them and by what rule, whether causes are removed, and how your own work is done. The top four rows of the ranking - organisation, changes, planned work, suppliers - are managed by discipline and agreements, not by budget. That is where most of the downtime lies that you can cut without buying anything.
You can build such a ranking from your data yourself, with the method from section 8, or with our help: send an export of incidents for six months or a year - we will return an analysis: where downtime concentrates, whether failures have a schedule, and which outages pull others along. What to do with each finding - in "What comes next".
Discuss the result with an operations engineer - .