Forecast and risk

How many more outages and when the next one

The main service was down again and everyone was "running around" once more. Nobody can say when it will end. The tech people promise something, of course, but after the third time few believe them. Sounds familiar?
Usually this raises worries about how many more such problems to expect and how stable the situation really is - is it time to prepare for bigger problems.

The task

Without a deep analysis, the problems of operations are not easy to find, but you can work out a couple of figures that make some points clear:

  1. How many incidents to expect in the coming week or month. On this basis you can already plan on-call duty, sprints and so on.
  2. How likely is one "really long" outage. Whether to prepare escalation and a reserve in advance (which usually turns into a whole budget), or it is a rarity whose risk is cheaper to ignore.

How it is calculated

Two methods from the family of forecasts - "what happens next?"

  • Forecast of event frequency (Poisson)

    Basis: the current statistics of the breakdown rate are used.

    Long-term changes: if systems began to fail more often (or less often) and the new pace holds for more than two months, the forecast is built only on the new data, and the old data is dropped.

    Short-term spikes: if the outage rate changed only recently, the main forecast does not change: in a couple of weeks it is impossible to tell chance from a stable worsening. But a backup scenario is also calculated: "what happens if the current spike becomes the norm".

    Instability: if failures come unevenly (none for a while, then many at once), the forecast gives a wider spread: not "there will be exactly 5 outages", but "there will be from 2 to 8".

    Forecast horizon: you cannot predict the future for a long time if there is little past data. The forecast period is no longer than the history collected: with two weeks of data, we predict only one week ahead.

  • The risk of long downtime

    The point: to estimate the probability of long system failures.

    Ordinary statistics: if we are interested in an outage length that has already occurred in the past, the estimate is built by simply counting such cases in the history.

    Unprecedented failures: if we need to estimate the risk of a failure longer than all past records, a special mathematical model for rare events is used. This is a theoretical calculation, so a range of probability is given, not an exact number of cases.

    Data filtering: only critical incidents (priorities 1 and 2) count as downtime. Non-urgent tickets that hang for a long time do not count as downtime: while they wait in the queue, the service itself keeps working.

    Result: what share of all incidents turns into a really long downtime, and how many such cases to expect per month.

What the forecast sees and what it does not

Based on the patterns and scenarios already found, these are the cases the report already handles:

What is in the logWhat the report says
Sudden spikes of failuresfor example, after a releaseThe system does not panic and does not raise the forecast several times over. The main calculation follows the old norm, and the spike gets its own scenario "what happens if this does not stop".
Seasonal overloadsfor example, DecemberA forecast for a quiet January is not built on a peak December. If there is data for last year, the system recognises this as seasonality.
Uneven outagesnow plenty, now noneIf failures come in waves, the system widens the forecast range to allow for this instability, and shows separately how much large one-off spikes spoil the picture.
Too little dataa history of only 2 weeksThe system refuses to guess a month ahead. The forecast is built for one week at most, with an explanation of the reasons.
Open incidentsThey count in the breakdown rate, but do not take part in the downtime calculation: they are not fixed yet. A separate summary is shown for them: how many are open and whether any are critical.
Auto-closing of ticketsthe system closes the ticket itself after 3 daysThe analyser recognises such artificial delays and does not count them as real long outages, so as not to distort the downtime statistics.
Mass clean-ups60 old tickets closed at onceThe system sees this anomaly, states the date of the "clean-up" and leaves these tickets out of the calculation of outage length.
Bureaucracy timeonly the time of formal closing is recordedIf only the closing time of the ticket is recorded, not the real recovery of the service, the report warns honestly about the error and asks you to record the exact repair date in future.

Example 1. How many outages to expect

Situation: an online shop. Two weeks ago a bad update was released, the number of errors rose sharply, and there was no rollback.

Next month - from 61 to 118 incidents (88 on average). If the current spike is not fixed - more than 300.

What the report shows: next month from 61 to 118 incidents are expected, 88 on average. But if the current spike is not fixed (now 10 outages a day instead of the usual 3), more than 300 outages will build up over the month.

The value: you get a realistic range, not a falsely exact figure. The load on tech support can be planned for the base level, without inflating the staff because of a temporary spike. And at the same time the price of delay is visible, if the bad release is not fixed right now.

Example 2. The risk of long downtime

Situation: the same online shop in normal operation. We estimate the risk that a critical service goes down for longer than 6 hours.

3% of all outages (every 33rd) last more than 6 hours - about 3 times a month.

What the report shows: 3% of all outages last more than 6 hours, on average about 3 times a month. If you take only the most critical failures (top priority), they drag on for 6+ hours about once in two months.

The value: it becomes clear that a long downtime is not an abstract scare but a regular event. The figure of 3% is misleading: it is the chance for one specific incident. Over a month, a long failure will happen almost for sure. So a plan approved in advance is needed for it: who decides on the switch to the reserve and who is responsible for communication with clients.

How to use this data in your work

Understanding the real frequency and length of outages moves IT management from "fire fighting" to planning:

  • On-call duty and hiringResources are assigned according to a mathematically based forecast, not to the emotions after the last failure. It is clear when a stronger team is needed and when basic support is enough.
  • Justifying the IT budgetReserving capacity is expensive. Now you have figures to prove to the business that buying "hardware" is necessary - or, on the contrary, to show that the risk of a major downtime is negligible and spending millions on a reserve makes no sense.
  • SLA managementYou stop promising clients and neighbouring departments an impossible availability of 99.99%, if the report shows regular multi-hour drops. The terms of contracts are adjusted to the real abilities of the system.
The same analysis on your data - "Run analysis" on the home page →

All data is processed in isolation and is not passed to third parties.