How the analysis works

Why one method is not enough: how we combine methods

You built a "top 10 systems by downtime" report, fixed the first line, and the failures are the same. Sounds familiar? One report answers one question. Where downtime builds up - yes. Why it builds up, when to expect the next outage and whether one breakdown pulls others along - no. So we do not look for one "smart" method, but combine several simple ones, in a certain order.

The main points in 30 seconds.

1. Each method sees its own thing and is blind to the rest. The worst system by downtime may turn out to be the victim of someone else's failure.

2. First the log is cleaned, then it is counted. False alarms, requests and duplicates change the numbers by a third.

3. The order matters: the forecast is built after the search for the date from which it got worse, and the list of the biggest failures after duplicates are merged.

4. The result is not a set of numbers, but an answer: what to fix first and why.

The task

A manager needs answers to a few simple questions. Each is answered by its own method:

  1. Where do we lose the most time? Which systems and which failures give the main downtime.
  2. How fast do we fix? The usual recovery time and how often long repairs happen.
  3. When does it break? At what hours and days, and are there outages that come back by the calendar.
  4. Did it get worse? From which date the outage rate changed and what to expect next.
  5. Why? Which breakdown pulls others along, which server breaks again and again, which metric warns in advance.

Diagram 1. In what order we calculate

Each step stands after what it depends on. To swap the steps is to get wrong numbers.

  1. One name - one system. Why first: if a cash register is written in three ways ("Cash register", "CASH REGISTER", "cash register " with a space), its downtime is split into three parts, and it drops out of the list of the worst.
  2. Remove what is not an outage. Service requests ("give access", "reset a password") and records "outage not confirmed". Why before the calculations: in a bank's failure log, 210 of 802 records turned out to be false alarms. With them the median repair is 14 minutes, without them 21.5. The repair looked a third faster than it is.
  3. Remove automatic closings. Tickets that a timer closed after three days or that were closed all at once when the queue was cleared. Why: otherwise the report will write "VPN - 3066 days of downtime", although the VPN was working.
  4. Merge one failure from many tickets. A communication link failure in a shop opens tickets for cash registers, Wi-Fi and payment. Why before the list of the biggest failures: otherwise one failure will take three places in it, and the real second biggest will drop out.
  5. Base numbers. How many outages, how fast we fix, where downtime builds up.
  6. When: first the week, then cycles. Why this way: the weekly rhythm ("every Thursday at night") is already found. The search for cycles does not repeat it and looks for what the week does not explain: "every 10 days", "at month end".
  7. First the date from which it got worse, then the forecast. Why this way: if there have been twice as many outages since August, a forecast from the yearly average will underestimate their number. If the new level holds for more than two months, the forecast is built only on it.
  8. Why: links, repeats on a site, a metric before an outage. Why last: they need a clean log and failures already merged. Otherwise a duplicate ticket would look like a "repeat after 2 minutes".

Diagram 2. Who does not see what and who covers it

Each method has a blind spot. We put next to it a method that covers it.

MethodWhat it seesWhat it does not seeWho covers it
Where downtime builds upWhich systems and failures give the main downtimeThat a system may suffer from someone else's failureLinks between outages
Recovery timeThe usual repair and long casesWhich exact failures pull the average upWhere downtime builds up (the list of the biggest failures)
Hours and days of the weekNight jobs, the load of a working dayA rhythm that is not weekly: "every 10 days", "month end"Cycles
CyclesAn outage that comes back by the calendarA one-off change: "it got worse from August"The date of the change in frequency
The date of the change in frequencyFrom which day there are more or fewer outagesThe cause of the changeYou: you check the date against your own work
Links between outagesWhich breakdown pulls others along within minutesA slow worsening over the hours before an outageA metric before an outage (needs an export of metrics)
Repeats on a siteA server or a shop that breaks again and againThat the trouble is in the whole system, not in one siteComparison with chance: the same outages handed out to sites at random
ForecastHow many outages to expect and how often a long downtime will comeNew causes that were not yet in the logRepeating the analysis once a quarter

Example. The worst system is not always the culprit

Situation: a retail chain with an online shop, a year of work, about 1100 tickets across 12 systems. Management asks what to fix first.

By downtime the worst is Wi-Fi in shops (23%). But a communication link failure pulls cash registers, Wi-Fi and payment after it.

One method (where downtime builds up): Wi-Fi in shops - 23% of downtime, cash registers - 16%, shop communication links - only third, 10%. The conclusion "let us change the Wi-Fi" suggests itself.

The second method (links): after a communication link failure, within 15 minutes tickets are opened for cash registers (10 times a year, by chance we would expect less than one), for Wi-Fi (9 times) and for payment (8 times). The link is the source that looks modest in the downtime ranking.

The other methods: there are no unusual hours, no cycles, and the rate did not change over the year (about 3 outages a day). So the matter is not in the schedule and not in a recent change.

The value: Wi-Fi stays the worst by itself too, and it must be looked into. But a backup link in the shops will remove part of the outages of three systems at once. Without the second method, this part of the downtime would be put down to cash registers, Wi-Fi and payment separately.

How to use this in your work

  • Do not fix by one rankingBefore you invest in the worst system, check whether it suffers from someone else's failures. Otherwise the money goes to the consequence.
  • First put the log in orderFalse alarms, requests and duplicates distort every number. The mark "outage not confirmed" and the ticket type in the export save hours of arguments.
  • Look at the answer, not at a set of numbersA good analysis ends with the phrase "fix this first, because...". If a report gives ten figures and not one such conclusion, the methods in it are not combined.
The same analysis on your data - "Run analysis" on the home page →

All data is processed in isolation and is not passed to third parties.