Checking changes

Did it get worse after a change

After a release or a migration it seems that there are more outages. The business people complain. The tech people see no problems (a real professional will always find a way to explain why it is not on their side)) Roll back? Or is it a couple of bad days and everything will settle by itself? In such a situation you need not even a referee, only numbers and facts.

The task

To understand from the incident log whether the outage rate changed over the period: when, by how much and what the trend of the new level is. "It seems" is a weak argument. A few bad days in a row happen without any cause, so the difference must be told from the usual swing.

How it is calculated

A method from the family of deviation search - "where and what is wrong?".

  • Automatic search for breaksThe algorithm itself goes through the whole history and compares the averages "before" and "after" each day. To rule out chance, it adjusts to the natural instability of your data: the risk of a false "finding" is only 3-5%.
  • Finding the nature of the changeThe system names not one point but an interval of dates in which the shift happened. If a sharp outage happened, the report records a "step". If failures grew slowly - a gradual rise.
  • Isolating temporary spikesAnomalies lasting from 5 days, where there are one and a half times more outages, are analysed separately. If everything returned to normal after them, the algorithm counts this as a one-off spike, and if there is a year of history, it compares with last year to check for seasonality.
  • Requirements for the dataYou need only the creation time of all tickets, open and closed. The search for changes starts with a history of 4 weeks, and on a year of data the algorithm reliably notices stable shifts from 20-25%.

What the method sees and what it does not

The threshold of this method is delicate: too low - and the method finds "worsening" in ordinary noise, too high - and it misses real ones. So we checked each picture on 40-150 artificial logs with an answer known in advance. Below is how often the method is right.

What was in the logWhat the report says
A sharp jump of outagesfor example, the contractor was changedA jump by a third was found in all checks. The report names a window of dates, not one day. It calls it a jump directly in three cases out of ten, and in the others it writes: "either a sharp jump or a gradual rise - the data cannot tell them apart".
A gradual rise of failuresfor example, new shops are added all the timeA rise of 30% over a year was found in 8 cases out of 10. The method never called it a sudden jump.
A small riseless than 20-25%It drowns in the usual swings from week to week. The report writes that no shift is visible, and says from what rise the method would have noticed it.
A temporary spike and a return to normalfor example, a bad release was rolled backA two-week spike was found in all checks, the dates are exact to the day. The main forecast is built on the usual level, not on the level of the spike.
A stable period without changesIn 95 cases out of 100 the report writes that there are no significant shifts. If outages come in waves, a false finding happens in 7 cases out of 100.
Worse in duration, not in numberthe same number of outages, but they take longer to fixThe method counts only the number of outages and does not see such worsening. Where to look - in the row "No changes visible" of the table below.
Too little dataan export for less than 4 weeksThe search for changes does not start: the history is too short to tell chance from a real shift.

Why the system cannot always tell a sharp jump from a smooth rise. In any company the number of incidents naturally swings from week to week, by 20-25% on average. If failures rose by a third, this rise drowns in the usual background "noise". So the exact date of the breakdown often cannot be set to the day: the system gives a "window of dates", and speaks of a sharp jump only when it clearly exceeds the usual margin of error.

The pictures and what to do about them

The ticket log rarely says why an outage happened (the "cause" field is often missing, or it holds what was clear to the shift at the start). So the report recognises typical mathematical scenarios and suggests where to look for the root of the problem and how to react.

ScenarioPossible causesHow to checkWhat to do
A sharp jump (a step)A big update was released, the contractor was changed, the number of points of sale grew sharply. Or the rules changed: the monitoring system began to create tickets itself.Pull the change history for these dates and look at the authors of the tickets: real people or monitoring robots.If the business grew - strengthen the on-call shift. If a release or a contractor is to blame - return the problem to its initiators. If new monitoring started working - there is no real worsening, only the counting rules changed.
A gradual riseA steady growth of load, ageing of equipment, build-up of technical debt.Compare the growth of the number of tickets with the growth of the number of users or shops by month.If there are as many tickets per point as before, this is the normal price of scaling. If there are more - look for a "bottleneck" in the architecture or worn hardware.
A temporary spike with a return to normalA bad release that was quickly rolled back, a failure at a provider, a sudden flow of customers because of a promotion.Study the update log and the ticket texts. If the texts are the same, this is not a hundred different breakdowns but one mass failure.Investigate the whole spike as a single incident: find the root cause and work out why the recovery took exactly this long.
A rise at the very end of the chartIn the first weeks it is impossible to tell what it is: a short spike, a seasonal load or a new, higher norm of failures.Compare with the same period of last year and check recent releases.Do not panic. The base forecast does not change, but an extra scenario "what if this lasts" is built. Final conclusions - in a month, when the statistics build up.
No changes visibleThe breakdown rate did not exceed the usual swings.-If the business still "feels it got worse", the reason for the discontent must be sought elsewhere. Most likely there are no more outages, but they take longer to fix, have become more critical for clients or hit the same sore service again and again.

An example of the result

Situation: a retail chain with an online shop. The system itself, without hinted dates, found an anomaly in the ticket history.

From 10 to 23 June 2025 - a temporary spike: about 8 outages a day instead of the usual 3. After 23 June the rate returned to normal.

The value:

  • We got exact limits for the investigation: it was enough to check the log of releases and changes for exactly these two weeks to find the cause.
  • The forecast for the next month is not distorted and is built on the usual level, about 3 outages a day: two bad weeks hardly move it.
  • If the export had broken off right in the middle of the spike, the system would not have drawn hasty conclusions and would have warned honestly that it is not yet clear: a short failure, seasonality or the start of a long worsening.
The same analysis on your data - "Run analysis" on the home page →

All data is processed in isolation and is not passed to third parties.