MTBF and MTTR: how often it fails and how fast it is fixed
MTBF (Mean Time Between Failures) - how long a service works from one outage to the next.
MTTR (Mean Time To Recovery) - how long it takes to bring a service back to a working state.
These are two basic measures of reliability. Together they show which services fail more often than others and where the company loses the most time on repairs.
The task
To calculate for each service how often it fails and how fast it is restored, and to find out which metrics can really be trusted in planning and in reports.
Two main mistakes in the calculation
- Describing recovery time with one average. A few long outages push the average up a lot, and it no longer shows how long an ordinary incident lasts.
- Counting the failure rate as a total across all services. A figure like "an outage every 8 hours" for the whole company does not show which exact service is unreliable.
How it is calculated
A method from the family of descriptive analytics.
- Recovery time (MTTR). For each incident we record the time from the start to full recovery. From this data we calculate three key figures:
Median - the point by which 50% of all incidents are closed.
Mean - the total duration of all outages divided by their number.
P95 (95th percentile) - the limit beyond which only 5% of the heaviest outages last.
We also calculate what share of the total downtime these 5% of heavy outages make up. If it is half or more, the recovery process in fact splits into two different modes: standard and emergency. - Time between failures (MTBF).
For one service: the observation period divided by the number of its failures.
For the whole company: the same approach shows the overall incident rate in the infrastructure.
What the data must contain:
The exact start and end time of each incident (or the final duration in minutes/hours).
A link from the outage to a specific service (to calculate MTBF).
Note: incidents still open at the time of the calculation are not included in MTTR.
An example of the result
An example from bank processing.
The main mass of small outages sits to the left of the median. To the right of the P95 line is a rare group of heavy outages: there are few of them, but they make up most of the total downtime.
Practical value
- Choosing metrics for reports. In reports and KPIs, show the median and P95, not the arithmetic mean.
Why: the mean (3 h 13 min) gives a false picture. Because of it, the business thinks that any ordinary outage is fixed in three hours (while half are solved in 22 minutes), and at the same time underestimates how heavy real outages are (9+ hours). - Separate rules and escalation. Outages split into fast (mass) and long (systemic) ones, so the way of working with them must differ. Escalation time limits work better when tied to the real distribution (median and P95):
Up to 30 minutes: the on-call shift fixes it by standard instructions (runbooks).
After 30-60 minutes (the standard time is exceeded): L3 engineers join and an Incident Commander is appointed.
After 2 hours: escalation to the head of the department and a response team is set up.
After 4-8 hours (close to P95): the crisis protocol starts (an architect and vendors are called in, work moves to backup schemes). - Where to put attention and resources. The main development and operations effort should go to the 5% heaviest outages. Cutting their number or duration gives the biggest effect, because they make up 72% of the total downtime.
- Prioritising services (by MTBF). Comparing time between failures gives a ranking of problem systems and sets the order of work:
Payment of services: an outage every 3.9 days (the most problematic service)
Processing center: an outage every 5.0 days
SBP transfers: an outage every 5.5 days
This list shows which service to start engineering improvements with, and serves as the starting point when you check the results after a quarter.