Links between outages

Which outages pull others along: how to find hidden links in the flow of incidents

In a shop the communication link is gone. Five minutes later three more tickets land in the Service Desk: the cash registers, the guest Wi-Fi and card payment have dropped. The likely result is that three different IT groups try to fix three "different" outages in parallel.
At the end of the month the manager sees in the report a rise in incidents and a fall in SLA for three services, although the root of the problem was one. From a flat ticket log it is impossible to tell where the root cause is and where the consequence is. While these links are hidden, the company spends engineering resources on removing symptoms, and in the reports one failure counts as several independent ones.

The task of the algorithm

The method of automatic search for links analyses raw exports from Service Desk systems and answers three questions:

  1. Which outage is the trigger? After the failure of which system do tickets for other services open like an avalanche.
  2. Is it a pattern or a coincidence? How many times a similar scenario repeated in practice, and how many such matches simple chance would give.
  3. How many real failures happened? What share of tickets is just an "echo" of a failure already under way, or duplicates that should be merged into one incident.

How the method works

The method belongs to the algorithms for finding associative links and is adapted to the specifics of incidents.

  • 1. The observation window (15 minutes)

    After any ticket appears, the algorithm watches a time window of 15 minutes: did tickets for other systems open in this time.

  • 2. Filtering random matches and noise

    The most frequent mistake in analytics is to take daytime activity for a relation. If two office systems break more often by day, it does not yet mean that one brings the other down.

    The base rhythm is taken into account: the algorithm compares matches with the usual load of the second system in these weeks, on this day of the week and in this hour.

    Leaving out a "storm day": days when everything fell at once (for example, the data centre lost power) are excluded from the search. In a storm everything matches everything, which creates false links.

    Check of stability: a link is accepted as real only if it repeated on at least 4 different days and holds when the two peak days are removed.

  • 3. Finding the direction: who is the main one

    A directed link (A → B): if system A almost always fails first, it is recorded as the root cause.

    A link "with no first" (A ↔ B): if sometimes one and sometimes the other system is first, the algorithm marks them as victims of a common cause (for example, a likely power failure or shared infrastructure).

  • 4. Merging incidents (deduplication)

    Continuation: a ticket for a linked system within 15 minutes is counted as an echo of the first outage, not as a new independent incident.

    Duplicates: tickets for the same system a few minutes apart are merged.

Important: a statistical link is not the same as a physical cause: A may not break B directly, and both may be brought down by a third system C. But the method shows exactly where to look for the root of the problem, and reduces information noise.

What the algorithm sees in real data and what it is protected from

Raw Service Desk logs are almost never perfect: they are full of repeated storms from monitoring, blurred time and human factors. This is how the method handles typical scenarios:

What happenedWhat is in the logHow the algorithm reactsPractical effect
A chain reactionThe communication link went down, then the cash registers, Wi-Fi and payment terminalsFinds the first system and builds the chain of consequencesIn the list of the biggest incidents the failure stands once instead of four times; the "culprit" is visible at once
A hidden common causeThe power went off; first by chance fall now the cash registers, now Wi-FiRecords a link "fail together, with no clear first"Tells engineers to look for an infrastructure cause (power, a switch)
Two consequences of one outageCash registers and Wi-Fi fell because of the communication linkDoes not link the cash registers and Wi-Fi directly to each otherRules out false guesses that the cash register software breaks Wi-Fi
Duplicates and "bounce"Monitoring or users raised several tickets for one outage within 10 minutesMerges repeats into one failure and counts their shareShows how many tickets are repeats and by how much the number of outages is inflated
A storm (mass outage)Because of a failure in the data centre or at a large provider, everything fails at onceExcludes such days from the search for linksProtects from hundreds of junk links "everything with everything"
Low quality of dataThe system is not given or the time is recorded only to the hourDoes not look for links and writes in the report whyProtects from unreliable conclusions when the records are kept carelessly

Reliability of the check: to build a margin of accuracy, besides tests on real databases, the algorithm was tested on hundreds of artificially generated logs with no real links - up to 30 systems and 5 500 tickets a year, including days of mass outages. A false link was found in only one year out of 200.

What to do if the algorithm found a link

When the model shows a regular link between outages, engineers get a ready pattern to work with. The table shows how to read the combinations that were found:

ScenarioLikely causeHow to check the guessWhat to do
System A falls first, then B, C and DA dependency in the architecture: B, C and D cannot work without ADo the sites (shop, central rack, branch) of a pair of tickets matchPriority for redundancy: investment in the reliability of the one system A removes outages in three others
Systems A and B fall together in random orderA common infrastructure cause (power, a shared switch, the data centre)Find a common hardware or logical node for the sites of the two ticketsLook for a single point of failure (SPOF): make redundant not the systems themselves, but the node they share
System A falls after itself several times in a rowMonitoring "bounce" or duplicate tickets from usersThe sources of the tickets (a robot or a person) and matching textsLess alert fatigue: suppressing repeated alerts protects engineers from burnout
Different teams take the same problem in parallelOrganisational separation: each department sees only its own symptomCompare registration times and sites of tickets of different teamsA single control centre: consequence tickets are linked to the main one, the fix is run from one point

Example. A communication link and three ticket queues

Source data: a retail chain, 12 IT systems, about 1 100 tickets a year. When the internet link in a shop breaks, the first support line gets tickets for the cash registers, Wi-Fi and payment terminals. Each ticket goes to its own service group.

A shop link outage pulls cash registers, Wi-Fi and payment after it: 10, 9 and 8 times a year - by chance it would match less than once. 56 tickets are continuations of failures already under way.

What the algorithm found: within 15 minutes after an outage of "Communication links", tickets for cash registers were opened - 10 times out of 106 (by chance about 0.8 times), for Wi-Fi - 9 times (about 0.4), for the payment gateway - 8 times (about 0.2). 56 tickets out of 1 116 a year are a direct continuation of failures already under way.

What this gives:

  1. Saving IT resources. Instead of three parallel investigations, one person is made responsible for the communication link. The hours of narrow specialists are not spent on symptoms.
  2. Transparent KPIs and SLAs. You can see that the number of incidents in the reports is inflated by about 5% only because of echoes and repeats. Reporting to management becomes honest.
  3. Justifying the return (ROI) for management. The budget for a second, backup communication link is easier to justify: "protecting one link removes three categories of outages at the registers and payment at once".

The bar - how many times after a link outage, within 15 minutes, a ticket for another system was opened; the grey mark - how many chance would give.

How to use the results in your work

The dependencies that were found change the work of the IT department on three levels - from everyday ticket handling to strategic decisions.

  • Processes (Incident Management)The principle "one failure - one main ticket": consequence tickets opened in the 15-minute window are linked to the parent ticket (Parent-Child). Result: one notice for the business, one engineer on duty and no confusion when different teams do the same work in parallel.
  • Architecture and budget (Problem Management)A "trigger" system that pulls a cascade of outages after it is the main candidate for redundancy and modernisation. Result: a clear return on IT projects. Strengthening one critical point removes a layer of secondary tickets and the load on the second and third support lines.
  • Management and KPIs (Executive Reporting)Management sees which part of the tickets is duplicates and "echoes" of failures, and the real picture of service availability. Result: decisions on purchases and redundancy are made on facts, not on an inflated number of tickets.

The main result

Searching for links between outages turns the Service Desk log from a passive "graveyard of tickets" into a tool for point optimisation of the IT infrastructure. You stop spending resources on removing symptoms and get a map of the hidden dependencies of your systems.

The same analysis on your data - "Run analysis" on the home page →

All data is processed in isolation and is not passed to third parties.