Downtime priorities

Which incidents cause the most downtime

Hundreds, even thousands of incidents in the tracker over a year are normal. There are no resources to fix everything at once. But where to start so that downtime falls fast?

The task

We have not yet met a team that manages to handle everything, always. And many people first want to fix the outages that happen most often. But the incidents that cause the most downtime are not always the frequent ones. Here it helps to have a way to separate the few important ones from the majority that can wait.

How it is calculated

A method from the family of descriptive analytics.

  • Pareto analysis (the 80/20 rule). We sort incidents by their contribution to downtime and add them up - to the point where the total reaches 80%. Next to it is the Gini coefficient, a number that measures concentration.

The method counts the frequency and the share of downtime, but not the severity of a single catastrophic outage. So it always comes with an estimate of the risk of long downtime (a separate analysis).

An example of the result

On real data (bank processing, 592 incidents over a year, unconfirmed outages removed) the engine calculated the contribution of each incident to the total downtime:

8.8% of incidents gave 80% of all downtime. That is 52 incidents out of 592. The Gini coefficient is 0.875 (a measure of inequality, where 0 means all incidents are equal and values near 1 mean almost all downtime comes from a few). This is a strong concentration.

The practical conclusion: working on these 52 incidents matters more than working on the whole mass at once. The other 91.2% may also need fixing, but they hardly move downtime, even if you close all of them. The saving of effort is clear and can give a good commercial effect.

The same analysis on your data - "Run analysis" on the home page →

All data is processed in isolation and is not passed to third parties.