Escalation matrix and alert noise: whom to wake and when
IT managers have two complaints that sound like different problems. First: "the on-call team is flooded with alerts, the engineer looks at the screen and no longer sees anything". Second: "I learned about a key service going down from an angry client, not from my own people". Sounds familiar? In fact it is one problem from two ends. Until the flow of monitoring signals is split into those someone must act on and all the rest, the escalation scheme stays on paper: it says whom to call, but not on which signal.
Below - how noise works, how to measure it and how to cut it, and then how to build, from a clean flow, an escalation matrix that works at three in the morning and not only in a meeting. It rests on the ITIL 4 practices "monitoring and event management" and "incident management", and on Google's SRE principles of on-call work.
1. Where noise comes from
Noise does not appear at once. It builds up over years, and almost every source of it was a sensible decision at its time. That is why cleaning the flow is hard: you argue not with a mistake but with someone's past common sense.
- Thresholds. After a big failure the alert threshold is lowered with a margin. Such failures do not happen again, but the signals stay. A year later nobody remembers why disk usage at 70% is a reason to call, but switching it off is scary too.
- Watching hardware, not the service. Monitoring is set up from the hardware list: every server, port, disk. The failure of one switch turns into dozens of messages about everything behind it being unavailable, instead of one about the cause.
- No silence during work. Planned maintenance and restarts after updates create a flow of "failure" messages. The team gets used to ignoring them - and one day ignores a real outage that happened at the same time as the work.
- Flapping. A link or a service goes down and comes back several times a minute. Each change is a pair of messages. There may be no incident at all, but there are already a dozen lines in the feed.
- Signals without an addressee. The message goes to a shared chat of fifty people, where nobody is responsible for it. Such a signal is not handled - it is only read.
- Legacy of systems and people who are gone. Checks for a service that no longer exists; rules of an engineer who left three years ago. Switching them off is scary: nobody knows what will break.
All these causes have one result, and it is not technical but human: alert fatigue. Google describes it directly in its SRE book: when there are too many signals, people start to skim and ignore them and miss the real one hidden in the noise. The on-call engineer stops reading the feed and picks out familiar pictures. This works until the first important message in an unfamiliar form. This is how noise turns from an inconvenience into a business risk.
2. How to measure noise
"Too many alerts" is a feeling, not a measure. The conversation changes when a few numbers for the last six months are on the table. They are counted from the export of monitoring events and of incidents from the ticket system, which you already have. No separate project is needed.
| Measure | How it is counted | Guide | What it tells you |
|---|---|---|---|
| Share of signals with no action | Signals closed without a single action, divided by all signals | For those that wake a person - close to zero: by the SRE principle, every such alert must require action | The main measure of noise. If it is high, you can skip the others for now: clean first |
| Signals per incident | Number of signals over a period, divided by the number of incidents | Single digits, not dozens (our guide) | Dozens mean that repeats and related events are not merged, not that the infrastructure is "busy" |
| Share of flapping | Signals that cleared by themselves before the on-call engineer could react | A noticeable share is a reason to add a time delay | The signal fired on a random spike, not on a stable state |
| Night wake-ups with no incident | Night signals that did not lead to an incident, divided by all night signals | Each one is a candidate for removal from the night route | The cost in people: an engineer woken for nothing works worse all next day |
| Incidents per shift | Incidents that one on-call engineer got, on average per shift | By Google's guide - no more than two per 12 hours: handling one, with the fix and the write-up, takes about 6 hours | More than that - the engineer only fights fires and has no time to remove causes |
| Signals without an addressee | Rules that go to a shared channel, not to a role or an on-call duty | Zero | A signal without an owner is not handled - it is only read |
Guides with no source given are reference points from our practice, not a norm. A telecom network and an accounting system have different normal values. The point is not to compare yourself with someone else's number but to count your own and look at it again in a quarter: the direction matters more than the absolute level.
3. Three classes of events and the one-addressee rule
ITIL divides events into three classes. Informational: a fact is recorded, no action is needed (a backup is done, a user logged in). Warning: a measure is getting close to a limit, action will be needed, but not now (the disk is 80% full). Exception: a breach, action is needed at once (the service is down).
The value of this split is that the three classes take different delivery paths. Informational - only to the log: nobody reads it in real time, and nobody should. Warning - to the work queue, handled during working hours. Exception - to a specific person, with a demand to react. When all three are mixed in one feed, the sorting that the system should have done is done by a tired engineer. This is the technical cause of noise.
Our rule for the procedure: an alert must have an addressee, an action and a deadline. If one is missing, it is not an alert but a log entry. It repeats the principle: every signal that wakes a person must require action and thinking; if a mechanical response is enough, it is a job for automation, not for the on-call engineer.
4. Seven steps to cut the flow
The order is not random. The first three steps give a visible result in a few days and cost almost nothing - it is easy to get the team to agree with them. The last ones need work with data and talks with service owners.
- Merging repeats. A repeated message about the same state updates the open event instead of creating a new one (deduplication). One failure - one line with a repeat counter.
- Silence during work. Suppressing alerts during maintenance, tied to the change calendar. The side effect is worth more than the main one: without an accurate calendar you cannot set up suppression, so the calendar appears.
- Time delay. A state must last a set time before it becomes a signal. This removes flapping without touching thresholds.
- Removing signals with no action. Take the rules sorted by number of firings and ask one question about each of the top ones: what does a person do on getting this? No answer - the rule goes to the log. It is not switched off, it just stops waking people.
- Thresholds from data, not round numbers. The threshold is set from how the measure actually behaved over several months. "90% disk" is a round number. "The level above which the measure was 2% of the time, and each time it ended in an outage" is a threshold.
- Linking by dependencies. If you know that twenty servers sit behind a switch, its failure gives one event about the cause, not twenty-one (correlation). You need at least a rough dependency map - this is the longest of the seven steps.
- Regular review. Once a quarter - half an hour on the noisiest rules and on rules that never fired. The first are noisy, the second are most likely broken and silent not because all is well. Google SRE advises reviewing alert statistics with management in the same way - once a quarter.
What not to do: start by switching off rules in bulk "to make it quiet". Such silence looks exactly like the silence of broken monitoring, and later you cannot tell them apart. Every removed rule must go either to the log or to a report that someone reads on a schedule.
5. Priority: impact and urgency
An escalation matrix starts not with people but with priority. In ITIL, incident priority is made of two independent things: impact (how many users and services are affected, and which) and urgency (how fast the consequences become serious). They are often confused: an outage for one person can be urgent - for example, at a till at rush hour - while a mass small inconvenience is not urgent. There are usually four or five levels; below is a typical version with four.
| Priority | Typical content | Who learns at once | Guide for time limits |
|---|---|---|---|
| P1 · critical | A key service is fully down or down for a large part of clients; there is no workaround | On-call engineer, shift lead, head of operations, service owner | Response - minutes. First message to clients - within the first hour. Then updates at regular intervals |
| P2 · high | The service works worse than usual or redundancy is lost: the next failure will be critical | On-call engineer, shift lead; the manager - in a summary | Response - within the shift. Recovery - the same day |
| P3 · medium | A partial failure, there is a workaround, work goes on | The responsible team, during working hours | Planned queue. Wakes nobody at night |
| P4 · low | An inconvenience, appearance, a single request | The team's queue | As the team has capacity |
The times in the table are typical guides, not a commitment. Real ones come from agreements with the business and from what the team can really sustain. A time limit written nicely but impossible to meet is worse than none: it teaches people to treat the procedure as decoration.
A rule that saves many night arguments (our practice): loss of redundancy is P2, not P3. The service works, the client notices nothing, and you want to leave it until morning. But any next failure is now critical: the safety margin that the redundancy was built for is already used up.
6. Escalation matrix: who, when, how fast
In ITIL there are two kinds of escalation, and they should not be mixed. Functional: the task moves to someone with more knowledge or rights - from the on-call engineer to the specialist engineer, then to the vendor or contractor. Its levels are usually called support lines: L1, L2, L3. Hierarchical: the question goes up the management line when decisions, money, priorities or talking to the outside world are needed. The first speeds up the fix, the second removes obstacles; one incident may need both. The labels M1 and M2 in the table are ours, for short.
| Level | Who | What they do | When it starts |
|---|---|---|---|
| L1functional | Shift on-call engineer, service desk | Takes the signal, sets the priority, applies the known workaround from the instruction, keeps the record | At once, on any event of the "exception" class |
| L2functional | Specialist engineer: network, servers, application, database | Diagnosis beyond the instruction, configuration changes, finding the cause | There is no instruction or it did not help; the L1 time limit for this priority has passed |
| L3functional | Architect, developers, hardware or software vendor, contractor | Defect analysis, the fix, a request to the supplier's support | The cause is inside the product or needs an architecture change |
| M1hierarchical | Head of operations, service owner | Sets priorities between incidents, allows non-standard actions, tells the business how the work is going | P1 - from the first minute; P2 - when the time limit passes |
| M2hierarchical | IT director, company management | Talks with clients and the press, decisions involving money and risk, talks with the supplier at their level | By the wake-up rules - section 7 |
Three signs tell a working matrix from one that exists only on paper. By the timer, not by how people feel. When the time limit passes, the next level joins by itself, without the inner fight "call or wait a bit more". Escalation is not a punishment. If raising a question means being asked "why couldn't you handle it", the engineer will wait until the last moment. SRE has a concept for this - the blameless review: you review events and rules, not people. Every role has a backup. A role without a second contact is one specific person, and while that person is on holiday, the escalation level disappears.
A line that is often forgotten is escalation to an external supplier. It has its own time limits, contract number, contact and confirmation steps. If the on-call engineer does not have this at hand at three in the morning, recovery time is decided not by technology but by how fast the right phone number is found.
7. When to wake a manager
This is the most debated question in any team, and a written rule should decide it, not the character of the on-call engineer. The rule must be applied in thirty seconds, without long thinking. Below are our four criteria: if any one fires, it is a reason to call at any time of day or night.
- Clients or money are affected. A service used from outside is down or distorts data. Not "may be affected", but there is a sign that it is already affected.
- A decision beyond the on-call engineer's rights is needed. Switch off some functions, roll back an update, switch the site, spend money, bring in a contractor outside the contract.
- You need to talk to the outside world. Clients, the regulator, the press, a big partner. Silence at this moment costs more than the failure itself, and the on-call engineer should not speak for the company.
- The forecast is worse than the fact. Everything works now, but events lead to a critical state: space is running out, a queue is growing, the backup has failed, the supplier's failure is spreading. Waking someone now is cheaper than in two hours.
The rule needs two safeguards. "Woke someone for nothing" is not the on-call engineer's mistake if a criterion fired: you review the rule, not the person. "Did not wake when it was needed" is always reviewed - but as a defect of the rule, not as misconduct. A team that was told off once for a night call will most likely not call next time.
It helps to agree in advance on the format of the call itself. Thirty seconds: what does not work, since when, whom it affects, what has been done, what you need from the other person. A person just woken up takes in structure better than a story, and the decision is made in the first minute, not the tenth.
8. What the data shows
Everything above can be checked with a normal export for 90 days or more: the event log from monitoring and the incident log from the ticket system. Not with a team survey: a survey shows ideas, an export shows behaviour.
- Broken escalation: a big gap between the start of the outage and the start of work. If an incident lasts two hours and the fix itself took fifteen minutes, they fix fast but take long to find who should fix it. DevOps measures this part with mean time to acknowledge (MTTA). It is a routing defect, not a skills defect.
- Noise: the ratio of signals to incidents and the share of signals closed without a single action. Both are counted mechanically.
- Formal priorities: if almost all incidents in the database have the same priority, there is no prioritisation - there is a field in a form.
- Night load: the daily profile and the share of night wake-ups that ended in a real incident. This is the cost of the current rules, measured in people woken up.
- Repeats: the same outage comes back for weeks. This means escalation fights again and again the fire it has already fought. This is no longer a question for the matrix but for problem management - the ITIL practice that finds and removes causes.
A short conclusion. An escalation matrix does not work on top of noise: it assumes that a signal reaching a person matters. So the order is the reverse of the usual one: first split the flow into "needs action" and "to the log", and only then write down levels, time limits and wake-up rules. Otherwise you get a procedure that is followed in a daytime meeting and not followed at three in the morning.
If you want to estimate where your team is now, there are two ways. Quick, in words - the maturity test in our bot: 7 questions, about 5 minutes, at the end a level and a few tips. Precise, from data - send an export of incidents: we will return an analysis: where downtime concentrates, how fast things are fixed and at what hours outages pile up.
If you are working on this inside the team and want to discuss the result with an operations engineer - .