A metric as an early warning

Which metric warns of an outage

The online bank went down at lunchtime, the second time in a week. At the review the engineer on duty opens the charts: three hours before the failure the database processor load was already creeping up. The charts were there, the alarm was not. Sounds familiar?
Monitoring collects hundreds of metrics, but nobody knows which of them really warns of an outage and which just makes noise. What is needed is not one more dashboard, but an answer from your own history: which metric, and how many hours before an outage, starts to behave unusually.

The task of the method

To compare the incident log with monitoring metrics for the same period and answer three questions:

  1. Which metric changes before outages? Processor, memory, traffic, disk - or the node stops sending data altogether.
  2. How long before the outage? 15 minutes, a few hours or a day.
  3. Is it a warning or a coincidence? How often the same thing happens on ordinary days when there is no outage.

What to send: the incident log and an export of metrics (Zabbix, Grafana, Prometheus) for the same period - from 14 days and from 8 outages. Both files in one upload.

How it is calculated

  • 1. Two files on one time axis

    The system brings together the time of tickets and metrics, takes the time zone into account and finds the monitoring node by the name of the site in the ticket.

  • 2. The "usual" for each hour

    The processor at 10 in the morning has one norm, at 3 at night another. The system remembers the usual level of each metric for each hour of the day, weekdays and weekends separately. Night backup is not counted as strange.

  • 3. Windows before an outage

    We look at what happened 15 minutes - 1 hour, 1-6 hours and 6-24 hours before the ticket: was the metric clearly higher or lower than usual.

  • 4. Comparison with ordinary days

    The same hours on other days when there was no outage. If on an ordinary day the processor is that high in 1 case out of 100, and before outages in half of the cases, it is not a coincidence.

  • 5. Protection from chance findings

    There are hundreds of metrics, and some of them will "guess" outages just by chance. The system corrects for the number of metrics and requires that the matches fall on different days.

Important: a metric that grows before an outage is not yet its cause. The report names possible causes and which data would tell them apart.

What the method sees in real data and what it is protected from

  • The alarm opened the ticket itselfMonitoring saw "processor above 90%" and a minute later created a ticket. This is one and the same event, not a warning: the system looks no closer than 15 minutes before the ticket, and it recognises tickets like "High CPU on db-01" as the echo of an alarm.
  • The ticket was raised later than the outage beganThe service went down at 12:00, the ticket was opened at 12:40 - and the metric "grew 40 minutes before the outage". A rise less than an hour before the ticket the report calls "an early sign or the start of the outage", not an early warning.
  • The node went silent at the moment of the outageThis is a sign of the outage itself. The report will say "the node went silent at the moment of the outage in 12 cases out of 15", but it will not call it a warning.
  • The metric grows after the outageAfter a restart the service works through the queue, and the load stays for a couple of hours. If the next outage happened in these hours, such a rise is not counted as an early warning.
  • A day of a mass failureOne bad day with twenty tickets does not create a finding: the matches must fall on different days.
  • Hourly data for a yearZabbix keeps a minute-by-minute history for a month, and after that hourly averages. On hourly data a lead of less than an hour is not visible, and the report says so.
  • Little dataLess than 14 days of common period or less than 8 outages - the method draws no conclusions and writes why.

Reliability of the check: the method was checked on artificial months and years of monitoring - 12 servers with three metrics each, a daily load rhythm, night backups, noise. A real early warning 3-6 hours ahead was found in 40 months out of 40, and with only 8 outages in 35 out of 40. In ordinary months there were no false findings at all, and with 105 metrics in 2 months out of 40.

What to do with the metric that was found

ScenarioPossible causesHow to checkWhat to do
The metric grows hours before an outage
(database processor 3-5 hours ahead)
A resource runs out under load; a heavy scheduled job; a growing number of requestsOpen the chart of the metric for the day before the last three outages and the job schedule for these hoursA warning to the engineer on duty at this level; sort out the job or the load within a Problem ticket
The metric falls before an outage
(free memory 2-6 hours ahead)
A memory leak, a hung process, no space leftHow the metric changes between service restartsFor now - a planned restart before the dangerous level; in essence - fix the leak
A rise less than an hour aheadMost likely this is already the start of an outage that was not noticed at onceDoes the ticket have the time the outage started or was foundStart recording the time of detection; set an alarm on this metric, to know before the users
A metric of a shared resource before outages of different systems
(the network core link)
The systems depend on one nodeThe dependency map: what these systems work throughMonitoring and a reserve for the shared node instead of repairing each system
The echo of an alarm
(the ticket was opened by monitoring on this same metric)
The threshold is already set-If the alarm fires too late - lower the threshold or add a duration

Example. An online bank database

Input: a bank, 12 online bank servers, three metrics per server (processor, free memory, incoming traffic) for a month at a 5-minute step and an incident log - about 130 outages in the same month.
Situation: the database server falls about once every day and a half. Each time the ticket is "Service unavailable", engineers restart the service, and all works until the next time.

What the report showed:

  • The processor of the database server was above the usual for that hour about 4.5 hours before the outage - in 11 cases out of 22. In the same hours of other days - in less than 1% of cases.
  • The other 35 metrics behaved as usual before outages.

The value of such a result:

  1. A warning in advance. The engineer on duty gets a signal a few hours before the likely outage, not a call from clients.
  2. Where to look for the cause. Not "everything", but what loads the database in these hours: a scheduled report, heavy queries, a growing number of users.
  3. An argument for a decision. "Before half of the outages the database processor was above the norm 4 hours ahead" is a clear argument for optimisation or a more powerful server.
  4. Checking the result. After a month, export the data again and see whether the rise before outages is gone.

The red line - in what share of outages the processor was above the usual that many hours before the ticket; the grey one - the same in the same hours of other days. Where the red line separates from the grey one, the metric warns of an outage.

How to use this data in your work

  • On-call dutyAn outage known a few hours ahead is planned work, not a night emergency. The shift has time to move the load or restart the service at a quiet moment.
  • Monitoring without noiseIt becomes visible which metrics really warn and which only wake the engineer on duty. Alarms on the first should be strengthened, on the second - weakened.
  • Budget and SLAThe argument "before half of the outages the database hit the processor limit" is clear to management without technical details. A forewarned outage can be prevented, which means less downtime counts against the availability promised to clients.
The same analysis on your data - "Run analysis" on the home page →

All data is processed in isolation and is not passed to third parties.