Which outages come back in a cycle: from intuition to a forecast
"The night exchange with the database fell again. It seems something like this already happened, and before that too..." In IT support this feeling of déjà vu is common. But to prove a pattern "by eye" is almost impossible: between these similar events in the log lie hundreds of other, random tickets.
The task
If the analysis by hours looks for a hard link to the weekly schedule, the analysis of cycles is predictive analytics. The method looks for a hidden rhythm of incidents that come back at equal intervals or are tied to the calendar. The algorithm scans the log automatically and answers three questions:
- Which outage comes back in a cycle. It finds the vulnerable system and its step - every how many days, or on which exact days of the month the breakdown happens.
- How far it is not chance. It counts how many times the incident matched the calculated cycle and compares this with the usual background rhythm of the system, leaving out random matches.
- When to expect the next failure. It calculates the date of the nearest outage in the cycle on the principle "if nothing changes" - or records that the chain broke and the cause has already been removed.
How the algorithm looks for cycles
The method rests on trying hypotheses and a strict check against chance. The algorithm does not just look for spikes, but finds a mathematically confirmed rhythm for each specific IT system. How it works step by step:
- Separating systems and night outages. The log is split by system (for example, 1C, the website, billing). The algorithm pays special attention to night outages (from 00:00 to 07:00). At night users sleep, so a spike of tickets at this time is most likely the trace of a failed script or job.
- Trying steps. The system checks all possible intervals: every 2 days, every 10 days and so on, covering periods up to a quarter of the whole log length. It separately checks ties to specific dates (for example, "the last three days of each month").
- Leaving out weeks. Steps of 7, 14 and 21 days are not searched for here. They are found by the analysis by days of the week and hours, so that the same incident is not repeated in the reports.
- Protection from random noise. To become a "cycle", an outage must fall into the rhythm at least 4 times and clearly exceed the usual background of the system. The algorithm compares each suspicious day with the norm. A busy season or a one-off series of failures will not become a cycle. The cut-off threshold is high: on random data the system simply stays silent and gives no false alarms.
- Forecast of the next date. If the maths confirms an active cycle, the algorithm names the date of the next expected incident on the principle "if nothing changes".
- Checking the fix. If there were no outages in the last three expected days of the cycle, no forecast is made. The report records that the cycle has stopped - most likely the engineers have already found and removed the cause.
Working with real data: what the algorithm finds
Incident logs are always full of "noise". The method is built so that it does not present wishes as facts. On ordinary logs without patterns, even with large one-off failures, the system will simply say: "No cycles".
- How the algorithm handles the features of logs
Daytime outages: a daytime cycle is harder to find - it drowns in the natural flow of user tickets. The algorithm will show it only if the outage happens in almost every cycle. If the rhythm is weak, the system prefers to stay silent.
Short history: if the export covers less than 6 weeks, the algorithm will not look for cycles. The maths needs at least 4 repeats to tell a schedule from chance.
Rare events: one ticket a month is 12 points a year among thousands of other incidents. To confirm such a cycle you need a history of several years; on a one-year slice the algorithm will treat it as chance.
No time of day: if the log has only dates without hours and minutes, the analysis still runs, but without a separate look at night outages.
Decision matrix: what a cycle that was found tells you
| Type of cycle | Possible cause | What the team does |
|---|---|---|
| Every few days, at the same hourfor example, every 10 days at 03:00 | An automatic job on a schedule: a heavy exchange, a recalculation, a data export | Check the scheduler. Move the job to another time, give it more resources or add a check script before the start |
| Every few days, but at different times | An accumulation effect: a full disk, a memory leak, a blocked queue, a cache or log that is never cleaned. A resource runs out in the same period | Set up monitoring of how full the resource is. Set up automatic cleaning or a preventive restart of the service on a schedule |
| The same days of the monthfor example, the 28th to the 30th | Closing of the financial month, regular reports, mass payments | Agree a calendar of heavy reports. Give extra database capacity in advance and put a specialist on call for these dates |
| Every 30, 60 or 90 days | Expiry: SSL certificates running out, password rotation, the end of licences or API access keys | Check issue and expiry dates. Introduce control of expiry dates or auto-renewal scripts for certificates |
Example 1. A script failure with a 10-day step
Situation: a retail chain, about 1100 tickets a year. The night exchange of 1C:ERP with the website fails from time to time. In the morning the warehouse sees unshipped orders, the ticket is closed by hand and forgotten until next time.
What the analysis showed: at 03:00 at night 1C:ERP fails every 10 days - an outage was recorded in 18 cycles out of 37, while the norm for this system is 2. The last case is 20 December, the next day of the cycle is 9 January.
Result: engineers stop putting out symptoms in the mornings and go to the task scheduler. Adjusting one background process removes about 18 failures a year and the morning downtime of the warehouse.
Each circle is a day of the cycle: red - the system failed on this day, empty - the day passed quietly. Ticks - outages on other days. The empty red circle after the end of the export is the next day of the cycle, if nothing changes.
Example 2. A month-closing problem
Situation: the accounting department complains about 1C freezes at the end of the month. The IT service thinks the problem is exaggerated, as it sees no clear proof in the general flow of tickets.
What the analysis showed: in the last three days of the month 1C:ERP fails in 11 months out of 12 - 18 outage days against a background norm of 6. The next risk window is 29-31 January.
Result: useless arguments give way to an exact calculation. For the month-closing dates the IT department gives extra server capacity in advance and assigns a specialist on call.
Summary: how to use this data in your work
A cyclic outage is an incident that can be predicted, and so prevented. Working with the cycle search algorithm gives three practical results:
- Targeted on-call dutyAn engineer is assigned strictly to the dates of the cycle and checks the health of the system the day before, instead of clearing the backlog the next morning.
- Arguments for management and technical debtA repeating outage becomes a clear line in the work plan with transparent maths: which system suffers, how many failures happen a year and what effect a fix will give.
- An automatic check of the resultThe method shows clearly when the problem is gone: as soon as the root cause is cleaned up, the report records that the cycle has stopped ("no outages in the last cycles").