Which incidents cause the most downtime
Hundreds, even thousands of incidents in the tracker over a year are normal. There are no resources to fix everything at once. But where to start so that downtime falls fast?
The task
We have not yet met a team that manages to handle everything, always. And many people first want to fix the outages that happen most often. But the incidents that cause the most downtime are not always the frequent ones. Here it helps to have a way to separate the few important ones from the majority that can wait.
How it is calculated
A method from the family of descriptive analytics.
- Pareto analysis (the 80/20 rule). We sort incidents by their contribution to downtime and add them up - to the point where the total reaches 80%. Next to it is the Gini coefficient, a number that measures concentration.
The method counts the frequency and the share of downtime, but not the severity of a single catastrophic outage. So it always comes with an estimate of the risk of long downtime (a separate analysis).
An example of the result
On real data (bank processing, 592 incidents over a year, unconfirmed outages removed) the engine calculated the contribution of each incident to the total downtime:
The practical conclusion: working on these 52 incidents matters more than working on the whole mass at once. The other 91.2% may also need fixing, but they hardly move downtime, even if you close all of them. The saving of effort is clear and can give a good commercial effect.