Which outages pull others along: how to find hidden links in the flow of incidents
In a shop the communication link is gone. Five minutes later three more tickets land in the Service Desk: the cash registers, the guest Wi-Fi and card payment have dropped. The likely result is that three different IT groups try to fix three "different" outages in parallel.
At the end of the month the manager sees in the report a rise in incidents and a fall in SLA for three services, although the root of the problem was one. From a flat ticket log it is impossible to tell where the root cause is and where the consequence is. While these links are hidden, the company spends engineering resources on removing symptoms, and in the reports one failure counts as several independent ones.
The task of the algorithm
The method of automatic search for links analyses raw exports from Service Desk systems and answers three questions:
- Which outage is the trigger? After the failure of which system do tickets for other services open like an avalanche.
- Is it a pattern or a coincidence? How many times a similar scenario repeated in practice, and how many such matches simple chance would give.
- How many real failures happened? What share of tickets is just an "echo" of a failure already under way, or duplicates that should be merged into one incident.
How the method works
The method belongs to the algorithms for finding associative links and is adapted to the specifics of incidents.
- 1. The observation window (15 minutes)
After any ticket appears, the algorithm watches a time window of 15 minutes: did tickets for other systems open in this time.
- 2. Filtering random matches and noise
The most frequent mistake in analytics is to take daytime activity for a relation. If two office systems break more often by day, it does not yet mean that one brings the other down.
The base rhythm is taken into account: the algorithm compares matches with the usual load of the second system in these weeks, on this day of the week and in this hour.
Leaving out a "storm day": days when everything fell at once (for example, the data centre lost power) are excluded from the search. In a storm everything matches everything, which creates false links.
Check of stability: a link is accepted as real only if it repeated on at least 4 different days and holds when the two peak days are removed.
- 3. Finding the direction: who is the main one
A directed link (A → B): if system A almost always fails first, it is recorded as the root cause.
A link "with no first" (A ↔ B): if sometimes one and sometimes the other system is first, the algorithm marks them as victims of a common cause (for example, a likely power failure or shared infrastructure).
- 4. Merging incidents (deduplication)
Continuation: a ticket for a linked system within 15 minutes is counted as an echo of the first outage, not as a new independent incident.
Duplicates: tickets for the same system a few minutes apart are merged.
Important: a statistical link is not the same as a physical cause: A may not break B directly, and both may be brought down by a third system C. But the method shows exactly where to look for the root of the problem, and reduces information noise.
What the algorithm sees in real data and what it is protected from
Raw Service Desk logs are almost never perfect: they are full of repeated storms from monitoring, blurred time and human factors. This is how the method handles typical scenarios:
| What happened | What is in the log | How the algorithm reacts | Practical effect |
|---|---|---|---|
| A chain reaction | The communication link went down, then the cash registers, Wi-Fi and payment terminals | Finds the first system and builds the chain of consequences | In the list of the biggest incidents the failure stands once instead of four times; the "culprit" is visible at once |
| A hidden common cause | The power went off; first by chance fall now the cash registers, now Wi-Fi | Records a link "fail together, with no clear first" | Tells engineers to look for an infrastructure cause (power, a switch) |
| Two consequences of one outage | Cash registers and Wi-Fi fell because of the communication link | Does not link the cash registers and Wi-Fi directly to each other | Rules out false guesses that the cash register software breaks Wi-Fi |
| Duplicates and "bounce" | Monitoring or users raised several tickets for one outage within 10 minutes | Merges repeats into one failure and counts their share | Shows how many tickets are repeats and by how much the number of outages is inflated |
| A storm (mass outage) | Because of a failure in the data centre or at a large provider, everything fails at once | Excludes such days from the search for links | Protects from hundreds of junk links "everything with everything" |
| Low quality of data | The system is not given or the time is recorded only to the hour | Does not look for links and writes in the report why | Protects from unreliable conclusions when the records are kept carelessly |
Reliability of the check: to build a margin of accuracy, besides tests on real databases, the algorithm was tested on hundreds of artificially generated logs with no real links - up to 30 systems and 5 500 tickets a year, including days of mass outages. A false link was found in only one year out of 200.
What to do if the algorithm found a link
When the model shows a regular link between outages, engineers get a ready pattern to work with. The table shows how to read the combinations that were found:
| Scenario | Likely cause | How to check the guess | What to do |
|---|---|---|---|
| System A falls first, then B, C and D | A dependency in the architecture: B, C and D cannot work without A | Do the sites (shop, central rack, branch) of a pair of tickets match | Priority for redundancy: investment in the reliability of the one system A removes outages in three others |
| Systems A and B fall together in random order | A common infrastructure cause (power, a shared switch, the data centre) | Find a common hardware or logical node for the sites of the two tickets | Look for a single point of failure (SPOF): make redundant not the systems themselves, but the node they share |
| System A falls after itself several times in a row | Monitoring "bounce" or duplicate tickets from users | The sources of the tickets (a robot or a person) and matching texts | Less alert fatigue: suppressing repeated alerts protects engineers from burnout |
| Different teams take the same problem in parallel | Organisational separation: each department sees only its own symptom | Compare registration times and sites of tickets of different teams | A single control centre: consequence tickets are linked to the main one, the fix is run from one point |
Example. A communication link and three ticket queues
Source data: a retail chain, 12 IT systems, about 1 100 tickets a year. When the internet link in a shop breaks, the first support line gets tickets for the cash registers, Wi-Fi and payment terminals. Each ticket goes to its own service group.
What the algorithm found: within 15 minutes after an outage of "Communication links", tickets for cash registers were opened - 10 times out of 106 (by chance about 0.8 times), for Wi-Fi - 9 times (about 0.4), for the payment gateway - 8 times (about 0.2). 56 tickets out of 1 116 a year are a direct continuation of failures already under way.
What this gives:
- Saving IT resources. Instead of three parallel investigations, one person is made responsible for the communication link. The hours of narrow specialists are not spent on symptoms.
- Transparent KPIs and SLAs. You can see that the number of incidents in the reports is inflated by about 5% only because of echoes and repeats. Reporting to management becomes honest.
- Justifying the return (ROI) for management. The budget for a second, backup communication link is easier to justify: "protecting one link removes three categories of outages at the registers and payment at once".
The bar - how many times after a link outage, within 15 minutes, a ticket for another system was opened; the grey mark - how many chance would give.
How to use the results in your work
The dependencies that were found change the work of the IT department on three levels - from everyday ticket handling to strategic decisions.
- Processes (Incident Management)The principle "one failure - one main ticket": consequence tickets opened in the 15-minute window are linked to the parent ticket (Parent-Child). Result: one notice for the business, one engineer on duty and no confusion when different teams do the same work in parallel.
- Architecture and budget (Problem Management)A "trigger" system that pulls a cascade of outages after it is the main candidate for redundancy and modernisation. Result: a clear return on IT projects. Strengthening one critical point removes a layer of secondary tickets and the load on the second and third support lines.
- Management and KPIs (Executive Reporting)Management sees which part of the tickets is duplicates and "echoes" of failures, and the real picture of service availability. Result: decisions on purchases and redundancy are made on facts, not on an inflated number of tickets.
The main result
Searching for links between outages turns the Service Desk log from a passive "graveyard of tickets" into a tool for point optimisation of the IT infrastructure. You stop spending resources on removing symptoms and get a map of the hidden dependencies of your systems.